A method for grouped fuzz testing of software
By calculating the similarity distance between seeds and grouping, combining adaptive value sharing and simulated annealing algorithm to optimize energy distribution and variation methods, the problem of unused seed similarity in the existing fuzz testing system is solved, and the efficiency and effect of fuzz testing are improved.
Patent Information
- Application Number
- CN202111578752.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-22
AI Technical Summary
Existing fuzz testing systems fail to effectively exploit similarities between seeds, resulting in inefficiency of fuzz testing, and existing methods such as AFL-HIER may lead to explosive growth of nodes in seed division.
By calculating the similarity distance between seeds, grouping using clustering methods, and using adaptive value sharing and simulated annealing algorithms for energy distribution and variation learning, combined with deterministic variation scheduling, the fuzz testing efficiency is improved.
It realizes more efficient use of seed similarity information, optimizes the fuzz testing process, and improves the efficiency and effectiveness of fuzz testing.
Smart Images

Figure CN114281690B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software security, and in particular to a method for performing grouped fuzz testing on software. Background Art
[0002] Software has currently been applied in all aspects of people's lives, so ensuring software security is very important. On the one hand, it is necessary to standardize software design and development, and on the other hand, it is necessary to timely discover and repair vulnerabilities existing in the software through testing. Automated fuzz testing is a currently popular method for discovering software vulnerabilities, which discovers problems existing in the software by providing various inputs to the program and simultaneously monitoring the software state.
[0003] Coverage-guided Fuzzing (CGF, or coverage-based fuzz testing), represented by AFL (American Fuzzy Lop), has become a very popular fuzz testing technology because of its extremely high efficiency in discovering software vulnerabilities, and has had a greater impact in both the industrial and academic fields. This application mainly focuses on CGF-based grey-box fuzz testing, and at the same time the proposed method can also be applied in hybrid fuzz testing systems that combine grey-box and white-box fuzz testing (such as Driller, QSYM).
[0004] Recently, the research published by Bohme et al. in FSE’20 shows that improving the efficiency of the fuzz testing system rather than adding more machines is more important for discovering vulnerabilities. Currently, there are two operations closely related to the efficiency of the CGF fuzz testing system: allocating appropriate energy to seeds (that is, determining how many inputs to generate from the seeds, also known as energy scheduling), and mutating the seeds in an appropriate way. However, existing fuzz testing systems usually determine how to perform such operations based on individual seeds or all seeds as a whole. The original AFL mainly allocates energy to seeds according to the execution time and coverage of the seeds. AFLFast (CCS‘16) models seed discovery as a Markov chain and attempts to allocate energy inversely proportional to the stationary distribution of individual seeds in the Markov chain. EcoFuzz (Security‘20) models seeds as a variant of the adversarial multi-armed bandit, but still allocates energy to each seed individually. On the other hand, in order to mutate seeds in an appropriate way, the original AFL uniformly selects among all mutation operations, while MOpt (Security’19) learns the optimal mutation operation selection probability distribution through the use of a customized Particle Swarm Optimization (PSO) algorithm, but it only learns one probability distribution to guide the mutation of all seeds.
[0005] The reason for considering seeds individually or as a whole in existing fuzzing systems is that they implicitly assume that seeds are independent of each other. This assumption is actually incorrect. For example, if a descendant seed is generated from another seed by small changes (e.g., a one-bit flip), it is obvious that mutating them can generate similar new seeds. The two seeds are very close in the input space. In addition, seeds that execute the same execution path should also be similar. For example, in fuzzing tcpdump, seeds corresponding to the same protocol are similar. In this case, the seeds are close in the program state space. In addition, similar seeds may have similar preferences for mutation methods, etc. Our experiments have also confirmed this. In existing fuzzers, this similarity between seeds is surprisingly overlooked. For example, even if two seeds have a single edge coverage difference, they are still considered completely different nodes in the Markov chain of AFLFast and different multi-armed bandit arms in EcoFuzz.
[0006] The present invention proposes a method of grouping using the similarity between seeds and using the grouping to improve the efficiency of fuzz testing. It is worth mentioning that this is different from a recent work, AFL-HIER (NDSS‘21). AFL-HIER organizes seeds into a tree rather than a flat group. For example, it divides seeds at the top layer according to function coverage, at the middle layer according to edge coverage, and leaf nodes are divided by Hamming distance. However, considering the exponential number of function combinations, for example, this design may lead to an explosive growth of tree nodes. In addition, the hierarchical division of seeds may not well reflect the similarity of seeds. For example, the new seed s has similar byte values to the seeds in subtree T1 but has a similar execution path to the seeds in subtree T2 (this is because it is mutated from the seeds in T1). Then, in AFL-HIER, the seed s will be assigned to subtree T2, even though its byte values are also very different from its sibling seeds. The flat grouping method of the method of the present application is more flexible. Since the seed s is far from the seeds in both T1 and T2, it is very likely to form a new group and perform fuzz testing according to its own preferences. Summary of the Invention
[0007] The present invention provides a method for grouped fuzz testing of software. Before grouping, the distance between seeds is calculated through the execution path of the seeds, the seed byte values, etc., and then clustering methods are used for grouping. After grouping, energy allocation based on fitness sharing is performed on the seeds, some optimal mutation methods are learned within each group to improve the efficiency of fuzz testing, and the number of seeds that perform deterministic mutation within each group is scheduled.
[0008] Specifically, the present invention includes aspects such as grouping seeds, optionally performing energy allocation after grouping, learning and using the optimal mutation method after grouping, and optionally performing deterministic mutation scheduling. The technical solutions adopted for them will be introduced separately below.
[0009] The solution for seed grouping is introduced as follows:
[0010] (1) First, define the distance between two seeds according to seed similarity. The seed distance should consider both the literal byte values of the seeds and the similarity of the execution paths, and may also include other similarities such as memory and variables during execution.
[0011] In addition, the distance calculation should be efficient because there are usually tens of thousands of seeds and many distances need to be calculated. Therefore, we do not directly measure the Hamming distance between the literal byte values of two seeds or the execution path bitmaps. Instead, we first calculate the LSH (Locality-Sensitive Hashing) of the byte values and execution path bitmaps of each seed, and then calculate the distance between seeds by calculating the Hamming distance between the LSHs. When calculating the LSH of the execution path, it is based on its execution bitmap, and when calculating the LSH, multiple bits are first combined into one bit so that the probability of a bit being 1 is 0.5 to increase its entropy value.
[0012] (2) Then, use a clustering algorithm to group the seeds in the queue. Clustering methods include but are not limited to K-means and DBSCAN. The above distance is used during clustering. The calculation of the mean (centroid) is based on the mean of the corresponding LSHs, that is, the majority vote of each bit of the LSH. The value of k is calculated according to an empirical formula. In addition, during the grouping process, it can be tried multiple times and the best result can be used.
[0013] Since the CGF fuzz testing system will continuously discover and add new seeds, new seeds need to be added to the nearest group here. And when a given number of new seeds are found or a given time interval is reached, overall re-grouping will be initiated.
[0014] Optionally, energy allocation after grouping can be performed, and the solution is as follows:
[0015] (1) First, still calculate the energy allocated to a certain seed according to attributes such as execution time, just like a general CGF fuzz testing system.
[0016] (2) Use the fitness sharing method to correct the allocated energy. Fitness sharing is often used in multimodal optimization to avoid getting stuck in local optima, and here the energy can be better distributed into different seed niches.
[0017] Learning and using the optimal mutation method after grouping, including learning the probability distribution of optimal mutation operation selection and the probability distribution of optimal mutation position selection. After discretizing the mutation positions into segments, similar to the mutation operation selection, their solutions are as follows:
[0018] (1) When learning the probability distribution of optimal mutation operation selection within a group, different mutation operations are represented by a probability value, and the sum of all mutation operation probabilities is 1. Initially, all probability values are the same. Using the simulated annealing algorithm, after obtaining a new probability distribution each time a change is made, calculate the fuzz testing efficiency of the new probability distribution. If the efficiency is higher, accept it; otherwise, accept it with a certain probability until the temperature drops to a certain threshold.
[0019] (2) Optionally, after the temperature is lower than a certain threshold, instead of stopping learning, enter a new phase (exploitation phase). At this time, the probability distribution will calculate new probabilities using the number of mutation operations selected when discovering new seeds and update them to the original probabilities.
[0020] (3) When learning the probability distribution of optimal mutation position selection within a group, divide the seeds into a certain number of regions (discretization) according to their positions, and each region is represented by a probability value. Subsequently, it is similar to learning the probability distribution of optimal mutation operation selection.
[0021] (4) When mutating the seeds within a group, apply the learned mutation information, including the probability distribution of optimal mutation operation selection and the probability distribution of optimal mutation position selection, to mutate the seeds.
[0022] Optionally, a deterministic mutation schedule within the group can be added to achieve a trade-off between the completeness and cost-effectiveness of testing. Its technical solution is as follows:
[0023] (1) The user provides the ratio of seeds for deterministic fuzz testing.
[0024] (2) During fuzz testing, when determining whether to perform deterministic fuzz testing on each seed in each group, it is decided according to whether each group has the same number of seeds that have completed the deterministic phase. Note that this is only done in the default mode, and if the user explicitly specifies to skip, the deterministic phase can also be completely skipped.
[0025] The beneficial effects of the present invention are as follows: By grouping seeds using the similarity between seeds, it is possible to more effectively learn the fuzz testing preference information of similar seeds and more efficiently guide fuzz testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic flowchart of fuzz testing seeds after grouping provided by an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram for comparing the exploration state space methods of fuzz testing before and after grouping in the embodiments of the present invention. Specific embodiments
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0029] The present invention first needs to group the seeds. In this embodiment, the grouping includes the following steps:
[0030] (1) First, define the distance between two seeds according to the seed similarity. s i and s j The distance between two seeds is defined as d(s i , s j ) = d v (s i , s j ) + d e (s i , s j ), where d v and d e respectively represent the byte value distance and the execution path distance.
[0031] In this embodiment, for the sake of computational efficiency, first calculate the 32-byte LSH of the byte value and the execution bitmap of each seed and splice them into a 64-byte LSH, which is stored in the data structure of the seed in memory. Then, calculate the distance between seeds by calculating the Hamming distance between the LSHs of two seeds. The length of the LSH is related to the discrimination accuracy between seeds and can be adjusted. In addition, since the bitmap corresponding to the execution path may be relatively sparse, multiple bits can be merged into one bit before calculating its corresponding LSH to increase the entropy value. After theoretical calculation, approximately B bits are merged into 1 bit, where B is the size of the execution bitmap and b is the number of 1 bits in the bitmap.
[0032] (2) Then, use the clustering algorithm to group the seeds in the queue. In this embodiment, the K-means clustering method is used. The above distance is used during clustering. The calculation of the mean (centroid) is based on the mean of the corresponding LSHs, that is, the majority vote of each bit of the LSH. In this embodiment, the value of k is calculated according to an empirically determined formula and is related to the current number of discovered seeds. In addition, during each grouping process, clustering can be tried multiple times to use the best result, that is, use the clustering result with the smallest average distance between each seed and its group centroid.
[0033] In this embodiment, new seeds need to be added to the group closest to them. At the same time, when a given number of new seeds are found or a given time interval is reached, the grouping is restarted.
[0034] After grouping the seeds in the present invention, the process of fuzz testing the seeds is as Figure 1 shown, including the following steps:
[0035] (1) Step S101, obtain the grouping information g of the seeds; in this embodiment, some grouping-related information such as the seeds in the group and the information learned within the group are saved in the grouping information g;
[0036] (2) Step S102, use the method of fitness sharing to allocate energy p to the seed. In this embodiment, first, still calculate the energy for a certain seed according to attributes such as execution time like a general CGF fuzz testing system, and set it as a certain seed s i The allocated energy is η(i), and then use the fitness sharing method to correct the allocated energy to p is the final energy allocation, where the calculation method of the niche count is is the average value of all seeds m i We adopt the Simple Exponential Smoothing Method to give higher weights to the nearest values, and sh(d ij ) is the sharing function in fitness sharing.
[0037] (3) Step S103, in this embodiment, based on the ratio of a seed for deterministic fuzz testing provided by the user, and the number of seeds in the current group and the number of seeds that have undergone deterministic testing, determine whether to perform deterministic fuzz testing on the seed to ensure that each group has the same number of seeds that have completed the deterministic stage.
[0038] (4) Step S104, loop p times: perform a series of operations of mutating the seed to generate an input and then providing it to the program for fuzz testing. In this embodiment, when mutating the seeds in the group, apply the learned optimal mutation operation selection probability distribution to mutate the seeds.
[0039] In this embodiment, when using the simulated annealing algorithm to learn the probability distribution of the optimal mutation operation selection within a group, the learning is carried out in two stages: the exploration stage and the exploitation stage. Different mutation operations are represented by a probability value, and the sum of the probabilities of all mutation operations is 1. At initialization, all probability values are the same. In the exploration stage, the simulated annealing algorithm is used. After obtaining a new probability distribution each time a change is made, the fuzz testing efficiency of the new probability distribution is calculated. If the efficiency is higher, it is accepted; otherwise, it is accepted with a certain probability until the temperature drops to a certain threshold. After the temperature is lower than this threshold, it enters the exploitation stage. At this time, the probability distribution will calculate a new probability using the number of operations selected when discovering new seeds and update it to the original probability. In this embodiment, simple exponential averaging is also used to update the probability.
[0040] Comparison between ordinary fuzz testing and the way of exploring the state space by fuzz testing after grouping seeds proposed in this application Figure 2 As shown. Before applying the grouped fuzz testing proposed in this application, a fuzz testing system like AFL independently uses each seed (black circle) to explore the state space. The energy assigned to each seed (the length of its arrow) and the method of mutating the seed (the direction of its arrow) are both determined separately. After applying the grouped fuzz testing proposed in this application, the state space is explored using seed groups. The energy assigned to each seed is related to the seed density near the seed, and the method of mutating each seed is learned and used by all the seeds within its group, achieving more efficient fuzz testing.
[0041] Those skilled in the art should understand that the embodiments in the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, the embodiments in this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0042] The embodiments in this application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products in the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.
[0043] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 or more processes and / or boxes Figure 1 or more boxes.
[0044] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 or more processes and / or boxes Figure 1 or more boxes.
[0045] Although the preferred embodiments in the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present application.
[0046] Obviously, those skilled in the art can make various changes and variations to the embodiments in the embodiments of the present application without departing from the spirit and scope of the embodiments in the embodiments of the present application. Thus, if these modifications and variations of the embodiments in the embodiments of the present application fall within the scope of the claims of the embodiments of the present application and their equivalent technologies, the embodiments of the present application are also intended to include these changes and variations.
Claims
1. A method for grouped fuzz testing of software, characterized in that, The method specifically includes the following steps: (1) Group the seeds; Calculate the distance between seeds according to the execution path of the seeds and the byte values of the seed content, use the distance between seeds to define the similarity between seeds, and group the seeds through the similarity between seeds. The number of groups is dynamically determined; The newly added seeds will be added to the group closest to it. When the number of added seeds reaches the preset threshold, all seeds will be regrouped as a whole. The grouping between seeds is carried out by means of a clustering method; learn the optimal mutation information of each group within the group, rather than learning the same optimal mutation information for all seeds; the optimal mutation information learned in each group includes the probability distribution of the selection of the optimal mutation operation and the probability distribution of the selection of the optimal mutation position; the mutation information learned in each group will be used to guide the mutation of the seeds within the group; (2) Obtain the grouping information g where the seeds are located; the number of groups is dynamically determined; (3) Use the method of fitness sharing to allocate energy p to this seed. Specifically: First, calculate the allocation energy η(i) of seed s according to the CGF fuzz testing system, and then use the fitness sharing method to allocate the fuzz testing energy of the seed. i where η(i) is the allocation energy, and then use the fitness sharing method to allocate the fuzz testing energy of the seed. where p(i) is the final energy allocation. is the number of niches, calculated by the distance between seeds. is the average value of the number of niches of all seeds, α is the weight value, and sh(d ij ) is the sharing function in the fitness sharing method. (4) Based on the ratio of a seed for deterministic fuzz testing provided by the user, and the number of seeds in the current group and the number of seeds that have undergone deterministic testing, when deciding whether to perform deterministic fuzz testing on the seed, schedule the number of seeds that perform deterministic mutation in each group to ensure that each group has the same number of seeds that have completed deterministic mutation; (5) Mutate the seeds to generate inputs and then provide them to the program for fuzz testing.
2. The method for grouping and fuzz testing software according to claim 1, wherein: The clustering method described includes but is not limited to K-means and DBSCAN.
3. The method for performing grouped fuzz testing on software according to claim 1, wherein: Learning the probability distribution of the selection of the optimal mutation operation and the probability distribution of the selection of the optimal mutation position within the group is carried out by means of a simulated annealing algorithm, and the learned optimal probability distributions are all discrete probability distributions.
4. A method for grouped fuzz testing of software as claimed in claim 1, wherein: When calculating the distance according to the execution path of the seeds and the byte values of the seeds, first use the locality-sensitive hashing method to hash the execution path and the seed byte value distribution, and then calculate the Hamming distance between the LSHs as the distance between the seeds; When calculating the LSH for the execution path, it is carried out according to its execution bitmap, and when calculating the LSH, multiple bits will be combined into one bit first, so that one bit has a probability of 0.5 of being 1 to increase its entropy value.
5. A method for grouped fuzz testing of software as claimed in claim 3, wherein: When learning the probability distribution of the selection of the optimal mutation operation within the group, different mutation operations are represented by a probability value, and the sum of the probabilities of all mutation operations is 1. All probability values are the same during initialization; Adopt the simulated annealing algorithm. After obtaining a new probability distribution each time a change is made, calculate the fuzz testing efficiency of the new probability distribution. If the efficiency is higher, accept it, otherwise accept it with a certain probability until the temperature drops to a certain threshold; When learning the probability distribution of the selection of the optimal mutation position within the group, the seeds are divided into a certain number of regions according to their positions. Each region is represented by a probability value, and the sum of the probabilities of all regions is 1. All probability values are the same during initialization, and the temperature is reduced to a certain threshold through the simulated annealing algorithm.
6. The method for grouped fuzz testing software according to claim 3 or 5, characterized in that, It also includes: When using the simulated annealing algorithm to learn the probability distribution, instead of stopping learning after the temperature drops below a certain threshold, it enters a new stage. At this time, the probability distribution will calculate the new probability using the number of mutation operations selected when discovering new seeds and update it to the original probability.
Citation Information
Patent Citations
A program path sensitive grey box testing method and device
CN109902024A
Text configuration file-oriented fuzzy test method and device
CN111913877A