An Integrated Fuzz Testing Method and System Based on Function Call Graph Partitioning

Through the test task division method based on function call graph, the global seed corpus is associated with the subgraph, which solves the problem of repeated test cases and inefficiency in parallel fuzz testing, and achieves more efficient integrated fuzz testing.

CN116820925BActive Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310500549.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-06-27
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

The lack of effective testing task division mechanisms in parallel fuzz testing in the prior art leads to inefficient repeated test cases and tests, especially when integrated fuzz testing is performed using multiple heterogeneous fuzzers.

Method used

The test task division method based on the function call graph is used to divide the function call graph of the target program into sub-graphs, and the global seed corpus is associated with the sub-graph to form a sub-seed corpus, which is assigned to each heterogeneous fuzzer for testing, and is rotated regularly.

Benefits of technology

By effectively dividing test tasks, repeated test cases are reduced, the efficiency and coverage of integrated fuzz testing are improved, and the ability to discover code defects and vulnerabilities is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116820925B_ABST
    Figure CN116820925B_ABST
Patent Text Reader

Abstract

The present invention discloses an integrated fuzz testing method and system based on function call graph partitioning. The present invention uses multiple heterogeneous fuzzers to perform parallelized integrated fuzz testing on a program, which is divided into an exploration stage and a development stage. In the exploration stage, a traditional integrated fuzz testing method is used to pre-test the target program to provide a seed corpus and branch coverage information for the development stage. The development stage includes three steps: First, the function call graph of the target program is partitioned into subgraphs, which are used as sub-tasks for fuzz testing; Then, each subgraph is associated with some seeds in the seed corpus to generate a sub-seed corpus; Finally, a central manager is used to allocate and schedule all sub-tasks, and synchronize the seeds and branch coverage bitmaps of each fuzzer during the testing process. The present invention can give full play to the advantages of integrated testing of multiple heterogeneous fuzzers, improve the software code test coverage rate, and discover more code defects / vulnerabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of program testing, and provides an integrated fuzz testing method and system based on function call graph partitioning. It is a method and system for performing parallelized integrated fuzz testing on a program using multiple heterogeneous fuzzers, and particularly relates to a test task partitioning method based on a function call graph and a parallel testing framework for heterogeneous fuzzers, so as to improve the test coverage rate of software code and discover more code defects / vulnerabilities in the integrated fuzz testing method and system. Background Art

[0002] Fuzz testing is an efficient and scalable automated program testing method, and has currently become the main software vulnerability detection technology. Its core is to automatically generate a large number of random test cases, then use these test cases as inputs to run the target program, and monitor the state of the target program during the running process, so as to find test cases that may cause the target program to crash.

[0003] In order to improve the efficiency of fuzz testing, multiple fuzz testing instances are usually used in parallel, and multiple CPU cores are used simultaneously to increase the number of test cases executed per unit time. Most grey-box fuzz testing methods based on coverage feedback adopt the default parallel mode of AFL: each fuzz testing instance uses the same seed scheduling strategy and mutation strategy to test the target program, and synchronizes the new seeds discovered by each fuzz testing instance regularly. This method lacks a test task partitioning mechanism and is prone to generating a large number of duplicate test cases.

[0004] In the field of parallel fuzz testing, <<PAFL:extend fuzzing optimizations of single mode to industrial parallel mode>> (PAFL) proposes a task partitioning method of coverage bitmap equalization. However, the coverage bitmap cannot reflect the logical relationship of each branch in the program, resulting in the set of branches assigned to each fuzz testing instance by this method being fragmented, which affects the parallel testing efficiency. <<Collabfuzz:a framework for collaborative fuzzing>> (AFLTeam) proposes a task partitioning method based on the function call graph of the target program. By partitioning the function call graph into multiple subgraphs and assigning them to each fuzz testing instance in turn, higher testing efficiency than PAFL can be obtained, but it is only applicable to parallelized testing with a single fuzzer.

[0005] Since different fuzzers are good at exploring different target program spaces, parallelizing multiple instances of heterogeneous fuzzers can achieve better results than parallelizing multiple instances of any single fuzzer. <<EnFuzz: Ensemble Fuzzing with Seed Synchronization among Diverse Fuzzers>> proposed the idea of ensemble fuzz testing that uses multiple heterogeneous fuzzers to perform parallel testing on the target program, but lacks a test task partitioning mechanism. It only realizes the cooperation between heterogeneous fuzzer instances by periodically synchronizing the newly discovered seeds, resulting in the convergence of the seed corpora of each fuzzer and affecting the overall testing effect.

[0006] To address this issue, the present invention proposes a test task partitioning method based on function call graph partitioning for ensemble fuzz testing. By mapping the partitioning of the function call graph to the partitioning of the seed corpus, it realizes the support for heterogeneous fuzzers and can effectively improve the efficiency of ensemble fuzz testing. Summary of the Invention

[0007] The present invention proposes an ensemble fuzz testing method and system based on function call graph partitioning. The method includes an exploration phase and a development phase. As Figure 1 shown, in the exploration phase, heterogeneous fuzzers are used to perform pre-testing on the target program in parallel to obtain an initial global seed corpus. In the development phase, first, a graph partitioning algorithm is used to partition the function call graph of the target program into a given number of subgraphs. Then, each subgraph is associated with a part of the seeds in the global seed corpus to form a sub-seed corpus corresponding to each subgraph. Finally, each sub-seed corpus is assigned to a fuzzer for testing and rotated regularly.

[0008] Specifically, the method steps of the exploration phase and the development phase of the present invention are as follows:

[0009] Ⅰ. Exploration Phase

[0010] In this phase, a group of heterogeneous fuzzers are used to perform pre-fuzz testing on the target program to provide an initial global seed corpus for the development phase. The specific steps are as follows:

[0011] First, select a group of existing mainstream fuzzers and create fuzzer instances respectively.

[0012] Then, create a monitoring program to start each heterogeneous fuzzer to test the target program; the monitoring program maintains a global seed corpus.

[0013] During the testing process, the monitoring program monitors the new seeds generated by different heterogeneous fuzzers. If a new seed is found to be generated by a heterogeneous fuzzer, the monitoring program determines whether this seed can bring new global coverage; if it can, it adds it to the global seed corpus. The monitoring program regularly synchronizes the seeds saved in the global seed corpus to each fuzzer.

[0014] After the time budget (recommended to be 2 hours) of the exploration stage is exhausted, the global seed corpus generated by the pre-testing of heterogeneous fuzzers is obtained and provided for use in the development stage.

[0015] Ⅱ. Development Stage

[0016] In the development stage, a task partitioning method based on the function call graph is used to coordinate the parallel testing work of each fuzzer, as Figure 2 shown. The development stage can be further divided into three steps: function call graph partitioning, sub-graph association and seed mapping, and integrated fuzz testing with task partitioning.

[0017] 2-1. Function Call Graph Partitioning

[0018] Use the seed execution paths in the global seed corpus of the previous stage to update the function call graph of the target program extracted statically, and use the graph partitioning algorithm to partition the function call graph into sub-graphs. The specific steps are as follows:

[0019] 2-1-1. Duplicate Removal of Seed Corpus

[0020] Remove duplicates from the seeds in the initial global seed corpus obtained in the previous stage, and then use these seeds to obtain the branch coverage information of each function in the target program. The work process is as follows:

[0021] 2-1-1-1 Read all the seeds in the initial global seed corpus;

[0022] 2-1-1-2 Use the seed duplicate removal program (it is recommended to use the afl-cmin tool for duplicate removal of the seed corpus provided by AFL) to remove duplicates from all the seeds collected in the initial global seed corpus, and use all the de-duplicated seeds to form the global seed corpus;

[0023] 2-1-1-3 Use all the seeds in the de-duplicated global seed corpus as inputs in turn to run the target program with coverage information collection enabled, and count and record the branch coverage information of each function in the target program.

[0024] 2-1-2. Function Call Graph Partitioning

[0025] Update the function call graph of the target program using the global seed corpus and simplify it. Then, use an existing efficient graph partitioning algorithm to partition the updated function call graph to obtain a large number of partitions of the function call graph. The workflow is as follows:

[0026] 2-1-2-1 Static extraction of the basic function call graph of the target program (it is recommended to use llvm for extraction during compilation);

[0027] 2-1-2-2 Run the target program with execution path feedback in turn using each seed in the deduplicated global seed corpus as input, and record the function-level execution paths of the target program when using each seed as input;

[0028] 2-1-2-3 Traverse the execution paths of each seed, and add the functions (nodes) and function call relationships (edges) that do not exist in the function call graph of the target program to the function call graph;

[0029] 2-1-2-4 Delete the duplicate edges in the function call graph of the target program, the nodes that do not contain any basic blocks, and the nodes that cannot be traversed from the main node;

[0030] 2-1-2-5 Assign weights to each node in the function call graph based on the number of branches it contains and the number of branches that have not been covered yet, and attach the weights to the attributes of each corresponding node in the function call graph;

[0031] 2-1-2-6 Use an existing efficient graph partitioning algorithm to partition the function call graph of the target program into a large number of partitions with close internal connections.

[0032] 2-1-3. Partition merging

[0033] Merge the large number of function call graph partitions obtained in step 2-1-2 according to the call relationships (edges) and weight values, and combine them into a given number of subgraphs as each subtask of the fuzz testing. These subgraphs are expected to have similar workloads and close internal connections. The workflow is as follows:

[0034] 2-1-3-1 Find the partition containing the main function, traverse each partition in the function call graph, and merge the partitions with close call relationships, that is, merge the partitions that cannot be directly called by the main function with the partitions that can call it;

[0035] 2-1-3-2 Sort the merged partitions in descending order according to the sum of the weights of all functions they contain; assign each partition to the subgraph (subtask) with the smallest total weight of the partitions that have been assigned at that time in turn, and mark the merge result on the function call graph in the form of a minimum spanning tree (it is recommended to add an attribute of the subtask number to which each node in the function call graph belongs)

[0036] 2-1-3-3 Combine the partitions into subgraphs with a given total weight sum and similar weights based on the weights and of all nodes in each partition. The number of subgraphs is equal to the number of heterogeneous fuzzers, and each subgraph represents a subtask.

[0037] 2-2. Subgraph Association and Seed Mapping

[0038] Map the partitioning of the function call graph that the fuzzer cannot recognize to the partitioning of the seed corpus that can guide the fuzzer testing, and generate sub-seed corpora corresponding to each subtask. The workflow is as follows:

[0039] 2-2-1 Sequentially use each seed in the global seed corpus as the input to the target program, and run the target program with execution path feedback;

[0040] 2-2-2 Traverse each function on the function-level execution path of the seed, and cumulatively add its weight value to the weight sum of the seed corresponding to the subtask to which the function belongs;

[0041] 2-2-3 Assign the seed to the sub-seed corpus corresponding to the subtask with the highest weight sum.

[0042] 2-3. Integrated Fuzz Testing with Task Partitioning

[0043] Assign the sub-seed corpora corresponding to the subtasks obtained in step 2-2 to each fuzzer and rotate them regularly. During the fuzz testing process, modify the new seed synchronization strategy of the fuzzer and add a branch coverage bitmap synchronization strategy to the fuzzer.

[0044] When initially assigning subtasks, assign the corresponding sub-seed corpora to each fuzzer in sequence according to the number order; after that, rotate the subtasks responsible for each fuzzer every fixed period of time. The workflow is as follows:

[0045] 2-3-1 When initially assigning subtasks, assign a fuzzer instance (actually, the sub-seed corpus corresponding to the subtask) used for this test to each subtask in a one-to-one correspondence between the subtask number and the fuzzer number, and record the fuzzer used for each subtask;

[0046] 2-3-2 Check every fixed period of time (recommended to be one hour) whether all subtasks have been tested by all fuzzers in parallel. If not, rotate the fuzzer used for each subtask and record it; otherwise, end this round of development stage.

[0047] 2-3-3 New Seed and Branch Coverage Bitmap Synchronization Strategy

[0048] 2-3-2-1 Generation and Synchronization of New Seeds

[0049] ① After the fuzzer discovers a new seed, it places it in the new seed buffer pool instead of its own seed queue;

[0050] ② The monitoring program monitors the new seed buffer pool. Once a new seed is found, according to the association relationship between its execution path and the function call graph and subgraph, it is assigned to a certain subtask. Then, the seed is sent to the sub-seed corpus corresponding to this subtask;

[0051] ③ The fuzzer synchronizes the sub-seed corpus corresponding to the subtask it is testing at a fixed interval to obtain new seeds from it.

[0052] Synchronization of 2-3-2-2 Branch Coverage Bitmap

[0053] ① Before starting fuzz testing, first generate a global branch coverage bitmap using the global seed corpus, and initialize the branch coverage bitmap of each fuzzer instance with this global branch coverage bitmap;

[0054] ② During fuzz testing, read the branch coverage bitmap of each fuzzer instance at a fixed interval (recommended to be 5 minutes), add the newly covered branches to the global branch coverage bitmap. Then, synchronize the global branch coverage bitmap to each fuzzer instance.

[0055] According to the above method, the present invention implements an integrated fuzz testing system based on function call graph partitioning. The core part of the system is as Figure 3 shown, including: a global manager and a fuzzer.

[0056] The global manager includes five major modules: a function call graph partitioning module, a seed mapping module, a task management module, a new seed partitioning module, and a branch coverage bitmap synchronization module. Among them, the function call graph partitioning module updates the function call graph of the target program extracted in advance using the execution paths of all seeds generated in the previous stage, simplifies it, and then uses a graph partitioning algorithm to divide it into subgraphs representing each subtask; the seed mapping module divides the global seed corpus into sub-seed corpora corresponding one-to-one with subtasks according to the association degree between the execution path of each seed in the global seed corpus and each subgraph; the task management module assigns subtasks (sub-seed corpora) to each parallel fuzzer instance, and rotates the fuzzer used by each subtask at a fixed interval; the new seed partitioning module monitors the new seeds generated by each fuzzer instance in real time and assigns them to the corresponding subtasks according to their execution paths; the branch coverage bitmap synchronization module maintains a global branch coverage bitmap and synchronizes it with the branch coverage bitmaps of each fuzzer instance regularly during fuzz testing.

[0057] Corresponding to the new seed partitioning module and the branch coverage bitmap synchronization module in the global management module, a seed synchronization policy and a branch coverage bitmap synchronization policy are respectively added at the fuzzer end to jointly ensure that the fuzzer always keeps testing the subtask part assigned to it.

[0058] The main advantages of the present invention are as follows:

[0059] For the first time, task partitioning in integrated fuzz testing is realized. Using the task partitioning method based on the function call graph, the state space of the target program is divided into multiple parts, enabling parallel heterogeneous fuzzer instances to respectively test different parts of the target program, reducing duplicate work and improving the overall testing efficiency.

[0060] The present invention has good scalability and is easy to expand more heterogeneous fuzzers (mainly for AFL-based fuzzers). Compared with the previous parallel fuzz testing with task partitioning, using the task partitioning method of the present invention does not require a specific fuzzer to be paired. Only the new seed generation and synchronization policy of the AFL-based fuzzer needs to be simply modified, and a branch coverage bitmap synchronization policy is added, then it can be integrated into the system of the present invention. Description of the Drawings

[0061] Figure 1 Is the overall process of the integrated fuzz testing method based on function call graph partitioning.

[0062] Figure 2 Is a schematic diagram of the development stage of the integrated fuzz testing method based on function call graph partitioning.

[0063] Figure 3 Is a framework diagram of the development stage of the integrated fuzz testing system based on function call graph partitioning.

[0064] Figure 4 Schematic diagram of function call graph partitioning. Detailed Embodiment

[0065] The following will specifically describe a possible implementation scheme of the present invention in combination with the drawings of the present invention. The following implementation schemes are only used to further illustrate the present invention and are not the only implementation means of the present invention.

[0066] From a high level, the integrated fuzz testing system using the task partitioning method based on the function call graph can be divided into two major steps: the exploration stage and the development stage.

[0067] Ⅰ. Exploration Stage

[0068] In the exploration stage, the target program is fuzz tested using the traditional integrated fuzz testing method to quickly cover the branches in the target program that are easily covered by fuzz testing, and provide an initial global seed corpus for the development stage. The specific steps are as follows:

[0069] 1-1. Set the number of fuzzers n to be run in parallel for this fuzz test, and select these n fuzzers based on experience (for example, use 4 fuzzers in parallel, and the parallel fuzzers are AFLFast, MOPT, QSYM and radamsa);

[0070] 1-2. Start an instance in Docker for each fuzzer selected in step 1-1;

[0071] 1-3. Use these fuzzers to perform fuzz testing on the target program for t hours (recommended 2 hours), and regularly synchronize new seeds discovered by each fuzzer during the fuzz testing process.

[0072] II. Development Phase

[0073] The development phase includes three major steps: function call graph partitioning, subgraph association and seed mapping, and integrated fuzz testing with task partitioning.

[0074] 2-1 Function call graph division

[0075] Function call graph partitioning can be further divided into three steps: seed corpus deduplication, function call graph partitioning, and partition merging.

[0076] 2-1-1 Deduplication of seed corpus

[0077] Collect all seeds generated in the previous stage and remove duplicates, then execute these seeds to obtain the coverage of each function in the target program. The specific steps are as follows:

[0078] 2-1-1-1 Read the seed corpus generated by the fuzz test in the previous stage;

[0079] 2-1-1-2 Use a deduplication program (it is recommended to use afl-cim that comes with AFL) to deduplicate all seeds, and put the deduplication results into the global seed corpus (folder);

[0080] 2-1-1-3 Run all seeds in the global seed corpus on the target program with coverage information statistics enabled, and use a coverage information collection program (such as gcov) to collect and record the branch coverage information of each function of the target program (it is recommended to generate a record for each function, including the function name, the number of branches it contains, and the number of branches that have not been covered).

[0081] 2-1-2 Function Call Graph Partitioning

[0082] Divide the function call graph of the target program into a large number of partitions with close internal logical relationships. The specific steps are as follows:

[0083] 2-1-2-1 Static extraction of the basic function call graph of the target program (it is recommended to extract using llvm during compilation)

[0084] 2-1-2-2 Run each seed in the global seed corpus sequentially on the target program that can feedback the execution path, and record the function-level execution path of each seed;

[0085] 2-1-2-3 Traverse the execution path of each seed, and check whether there are functions or function call relationships that do not exist in the function call graph. If so, update the function call graph (the result is as shown in (b) below); Figure 4 as shown in (b) below;

[0086] 2-1-2-4 Use the coverage information of the target program obtained in step 2-1-1 to attach the number of branches included and the number of branches that have not been covered yet to each function node in the function call graph;

[0087] 2-1-2-5 Simplify the function call graph, delete the duplicate function call relationships, functions that do not contain any branches, and functions that cannot be reached from the main function (the result is as shown in (c) below); Figure 4 as shown in (c) below;

[0088] 2-1-2-6 Convert the function call graph into the form of a minimum spanning tree (the result is as shown in (e) below); Figure 4 as shown in (e) below;

[0089] 2-1-2-7 Calculate the weight value of each function node according to the number of branches included and the number of branches that have not been covered yet (the recommended calculation formula is as described later) and attach it to the attributes of the node;

[0090] 2-1-2-8 Use the tree partitioning algorithm (recommended lukes algorithm) to partition the function call graph in the form of a minimum spanning tree into a large number of partitions with strong internal correlation (the result is as shown in (d) below). Figure 4 as shown in (d) below;

[0091] Among them, the formula for calculating the function weight in step 2-1-2-7 is recommended as:

[0092]

[0093] Among them, B t represents the total number of branches included in the current function, B uc represents the number of branches included in this function that have not been covered by fuzz testing, and α, β, and γ are all constants (the recommended values are: α = 1, β = 3, γ = 3).

[0094] For the function nodes with weights exceeding the set threshold, reduce their weights. The weight value of each function in the final target program is calculated as follows:

[0095]

[0096] Among them, W t represents the total weight of all function nodes in the function call graph, and n represents the number of required subtasks.

[0097] The pseudocode of the above process is as follows:

[0098]

[0099]

[0100] As shown in the above pseudocode, first, use the execution paths of the seeds in the seed corpus to add function call relationships that cannot be obtained during static analysis to the function call graph (lines 3 to 15); then attach the number of basic blocks, the number of branches, and the number of explored branches contained in each function to each node of the function call graph (lines 17 to 20); delete the duplicate edges among them (line 22), and the function nodes that do not contain basic blocks or cannot be reached from the main function (line 23); convert the function call graph into the form of a minimum spanning tree (line 24); calculate the weight of each function node therein (line 26), and attach it to each function node of the graph (line 27); use the Lukes algorithm to divide the function call graph into subgraph partitions (line 29).

[0101] 2-1-3 Partition merging

[0102] Merge the large number of partitions obtained in step 2-1-2 into a given number of subgraphs (subtasks). The specific steps are as follows:

[0103] 2-1-3-1 Find the partition containing the main function;

[0104] 2-1-3-2 Traverse the other partitions, and merge the partitions that cannot be directly called by the main function with the partitions that can call it;

[0105] 2-1-3-3 Sort the merged partitions in descending order of the sum of the weights of all functions they contain;

[0106] 2-1-3-4 Sequentially divide each partition into the subgraph (subtask) with the smallest sum of weights of the partitions that have been assigned at that time;

[0107] 2-1-3-5 Mark the merge result on the function call graph in the form of a minimum spanning tree (it is recommended to add an attribute of the sub-task number to which each node of the function call graph belongs).

[0108] The pseudocode of the above process is as follows:

[0109]

[0110]

[0111] As shown in the above pseudocode, first find the MST partition where the main function is located (line 2); merge the partitions that cannot be directly called by the partition where the main function is located and the partition that can call this partition among the remaining partitions (lines 4 to 8); sort the merged partitions in descending order according to the sum of the weights of all functions they contain (line 10); assign them to the subtasks with the smallest total weight at that time in turn (lines 17 to 21); attach the subtask number to which each function node belongs to the nodes of the MST (lines 23 to 27).

[0112] 2-2 Subgraph Association and Seed Mapping

[0113] Divide the global seed corpus into sub-seed corpora corresponding to each subgraph (subtask) according to the sum of the weights of the functions belonging to each subgraph on the seed execution path. The specific steps are as follows:

[0114] 2-2-1 Run each seed in the global seed corpus on the target program that can feedback the execution path in turn;

[0115] 2-2-2 For the current seed, traverse each function on its function-level execution path and find the corresponding node in the function call graph;

[0116] 2-2-3 Accumulate the weight value of the current function (read from the node attributes) into the total weight of the subtask corresponding to this function (read from the node attributes) to which the seed belongs;

[0117] 2-2-4 Assign the current seed to the sub-seed corpus corresponding to the subtask with the highest weight sum.

[0118] In particular, to prevent the sub-seed corpus corresponding to a certain subtask from being empty, it is recommended to use a two-dimensional array to record the k seeds with the largest weight sum for each subtask when dividing the seed corpus (k is dynamic. When the total number of seeds is less than 100, k takes the integer part of the total number of seeds divided by 10, otherwise k takes 10). After all the seeds in the seed corpus are assigned, these seeds will be additionally copied to the corresponding sub-seed corpus.

[0119] The pseudocode for the above process is as follows:

[0120]

[0121]

[0122] As shown in the above pseudocode, each seed collected in the previous stage is executed in sequence (line 5); the total weight of the functions corresponding to each subtask on its execution path is counted separately (lines 9 to 15); it is assigned to the subtask with the highest corresponding weight (lines 17 and 18); to prevent the situation where a subset of a seed corpus is empty, the k seeds with the largest weight for each subtask will be additionally assigned to that subtask (lines 25 to 29), even if the weight of the seed for other subtasks is greater.

[0123] 2-3 Integrated Fuzzing with Task Partitioning

[0124] Use the subtask assignment and scheduling strategy to assign sub-seed corpora to each fuzzer and rotate them regularly. The specific steps are as follows:

[0125] 2-3-1 Assign the sub-seed corpora corresponding to each subtask to the fuzzer instances with corresponding numbers in sequence according to the numbers, and record the assignment results in the subtask management table;

[0126] 2-3-2 Every m hours (recommended to be 1 hour), check whether each subtask has been tested by all parallel fuzzers (check the subtask management table). If not, assign the word-seed corpus corresponding to each subtask to the fuzzer assigned to the next subtask in this test (specifically, the fuzzer used by the subtask with the largest number is assigned to the subtask numbered 1 in this test) and record it. Otherwise, end this round of development stage.

[0127] Additionally, it is recommended to save the newly discovered seeds to a cache pool (recommended to be implemented directly with a folder) instead of its own seed queue by modifying the save_if_instering function of the fuzzer. Modify the sync_fuzzers function to change the seed folder to be synchronized from the seed queues of other fuzzers to the sub-seed folder corresponding to the subtask currently being tested by the fuzzer instance.

[0128] For branch coverage bitmap synchronization, it is recommended to use named shared memory to save the branch coverage bitmap of each fuzzer instance, and use socket communication to send the shared memory name to the synchronization program. The synchronization program directly uses the method of reading and writing the shared memory (protected by the semaphore mechanism) each time to quickly synchronize the branch coverage bitmap.

[0129] To verify the effectiveness of the present invention, the method of the present invention is compared with the work AFLTeam that uses a single fuzzer with task partitioning for parallel testing and the original integrated fuzz testing without using task partitioning. Specifically, we allocated 4 CPU cores to each tool. Among them, both the present invention and the original integrated fuzz testing use four fuzzers, namely AFLFast, MOPT, QSYM, and radamsa, while AFLTeam uses 4 processes of its specially made fuzzer. We tested 6 real-world software for 10 hours and repeated the test 3 times. The final results are shown in the following table:

[0130]

[0131] As shown in the above table, the method of the present invention can cover more branches than the work AFLTeam that uses a single fuzzer with task partitioning for parallel testing and the original integrated fuzz testing without using task partitioning. Therefore, the method described in the present invention can effectively improve the overall efficiency of integrated fuzz testing.

[0132] It should be noted that the above implementation is only used to illustrate the present invention rather than to limit the present invention. The various specific numerical values used in the examples only serve as examples. And those skilled in the art can design other implementation manners to replace the above specific implementation manners without departing from the scope of the appended claims, and it should not be regarded as a new method. All equivalent changes and modifications made according to the scope of the invention claimed in the present application fall within the scope of protection of the present invention.

Claims

1. An integrated fuzz testing method based on function call graph partitioning, characterized in that: It is divided into an exploration stage and a development stage; In the exploration stage, the integrated fuzz testing method is used. First, a group of existing mainstream fuzzers are selected to create fuzzer instances respectively; Second, a monitoring program is created to start different heterogeneous fuzzers to test the target program. During the testing process, the monitoring program monitors the new seeds generated by different heterogeneous fuzzers. If it is found that a heterogeneous fuzzer generates a new seed, the monitoring program determines whether the new seed can bring new global coverage; If so, the new seed is added to the global seed corpus; The monitoring program regularly synchronizes the seeds saved in the global seed corpus to each fuzzer; The development stage includes three steps: function call graph partitioning, subgraph association and seed mapping, and integrated fuzz testing with task partitioning; 2-1 Function call graph partitioning: Use the seed execution paths in the global seed corpus of the previous stage to update the function call graph of the statically extracted target program, and use the graph partitioning algorithm to partition the function call graph into subgraphs; 2-2 Subgraph association and seed mapping: Map the partitioning of the function call graph that cannot be recognized by the fuzzer to the partitioning of the seed corpus that can guide the fuzzer testing, and generate sub-seed corpora corresponding to each subtask; 2-3 Integrated fuzz testing with task partitioning: Assign the sub-seed corpora corresponding to the subtasks obtained in step 2-2 to each fuzzer and rotate them regularly; During the fuzz testing process, modify the new seed synchronization strategy of the fuzzer, and add a branch coverage bitmap synchronization strategy to the fuzzer.

2. The integrated fuzz testing method based on function call graph partitioning according to claim 1, characterized in that The specific implementation of the function call graph partitioning described in step 2-1 is as follows: 2-1-1 Seed corpus deduplication, deduplicate the seeds in the initial global seed corpus obtained in the previous stage, and then use these seeds to obtain the branch coverage information of each function in the target program; 2-1-2 Function call graph partitioning, use the global seed corpus to update the function call graph of the target program and simplify it; Then, use the existing efficient graph partitioning algorithm to partition the updated function call graph to obtain a large number of partitions of the function call graph; 2-1-3 Partition merging, merge the large number of function call graph partitions obtained in step 2-1-2 according to the call relationship and weight value, and combine them into a given number of subgraphs as each subtask of the fuzz testing. The expected workload of this subgraph is similar and the internal connection is tight, serving as each subtask of the fuzz testing.

3. An integrated fuzz testing method based on function call graph partitioning according to claim 2, characterized in that The specific implementation of step 2-1-1 is as follows: 2-1-1-1 Read all the seeds in the initial global seed corpus; 2-1-1-2 Use the seed deduplication program to deduplicate all the seeds collected in the initial global seed corpus, and use all the deduplicated seeds to form the global seed corpus; 2-1-1-3 Input all the seeds in the deduplicated global seed corpus into the target program with the coverage information collection enabled, then run the target program, and count and record the branch coverage information of each function in the target program.

4. An integrated fuzz testing method based on function call graph partitioning according to claim 3, characterized in that The specific implementation of step 2-1-2 is as follows: 2-1-2-1 Static extraction of the basic function call graph of the target program; 2-1-2-2 Run the target program with each seed in the deduplicated global seed corpus in sequence, and record the function-level execution paths of each seed in the target program. 2-1-2-3 Traverse the execution paths of each seed, and add the functions and function call relationships that do not exist in the function call graph of the target program to the function call graph. 2-1-2-4 Delete the duplicate edges, nodes that do not contain any basic blocks, and nodes that cannot be traversed from the main node in the function call graph of the target program. 2-1-2-5 Assign weights to each node in the function call graph based on the number of branches it contains and the number of branches that have not been covered yet, and attach the weights to the attributes of each corresponding node in the function call graph. 2-1-2-6 Use an existing efficient graph partitioning algorithm to partition the function call graph of the target program into a large number of highly interconnected partitions.

5. The integrated fuzz testing method based on function call graph partitioning according to claim 4, wherein The formula for calculating the weight is: Among them, B t represents the total number of branches included in the current function, and B uc represents the number of branches included in this function that have not been covered by fuzz testing. α, β, and γ are all constants; For function nodes whose weights exceed the set threshold, reduce their weights. Finally, the weight values of each function in the target program are calculated as follows: Among them, W t represents the sum of the weights of all function nodes in the function call graph, and n represents the number of required subtasks.

6. The integrated fuzz testing method based on function call graph partitioning according to claim 5, wherein Step 2-1-3 is specifically implemented as follows: 2-1-3-1 Find the partition containing the main function, traverse each partition in the function call graph, and merge the partitions with close call relationships, that is, merge the partitions that cannot be directly called by the main function with the partitions that can call it. 2-1-3-2 Sort the merged partitions in descending order according to the sum of the weights of all functions they contain; sequentially assign each partition to the subgraph with the smallest total weight of the partitions that have been assigned at that time, mark the merge result on the function call graph in the form of a minimum spanning tree, and add an attribute of the sub-task number to which each node in the function call graph belongs. 2-1-3-3 Combine the partitions into a given number of subgraphs with similar total weights according to the sum of the weights of all nodes in each partition. The number of subgraphs is equal to the number of heterogeneous fuzzers, and each subgraph represents a sub-task.

7. An integrated fuzz testing method based on function call graph partitioning according to claim 6, characterized in that Step 2-2 Subgraph association and seed mapping are specifically implemented as follows: 2-2-1 Run the target program with each seed in the global seed corpus in sequence, and run the target program with execution path feedback. 2-2-2 Traverse each function on the function-level execution path of the seed, and cumulatively add its weight value to the weight sum of the sub-task to which the seed belongs corresponding to this function. 2-2-3 Assign the seed to the sub-seed corpus corresponding to the sub-task with the highest weight sum.

8. An integrated fuzz testing method based on function call graph partitioning according to claim 7, characterized in that Step 2-3 Integrated fuzz testing with task partitioning is specifically implemented as follows: 2-3-1 When initially assigning sub-tasks, assign a fuzzer instance used for this test to each sub-task in a one-to-one correspondence between the sub-task number and the fuzzer number. Actually, the sub-seed corpus corresponding to the sub-task is assigned, and record the fuzzer used by each sub-task. 2-3-2 Check every fixed period whether all sub-tasks have been tested by all fuzzers in parallel; if not, rotate the fuzzer used by each sub-task and record it; otherwise, end this round of development stage.

9. An integrated fuzz testing method based on function call graph partitioning according to claim 7 or 8, characterized in that The generation and synchronization of new seeds in Step 2-3-2 are specifically implemented as follows: ① After the fuzzer discovers a new seed, it places it in the new seed buffer pool instead of its own seed queue; ② The monitoring program monitors the new seed buffer pool. Once it discovers a new seed, according to the association relationship between its execution path, function call graph and subgraph, it divides the new seed into a certain subtask; then, it sends the seed to the sub-seed corpus corresponding to this subtask; ③ The fuzzer synchronizes the sub-seed corpus corresponding to the subtask it is testing at a fixed interval to obtain new seeds in it; The synchronization of the branch coverage bitmap in Step 2-3-2 is specifically implemented as follows: ① Before starting the fuzz testing, first generate a global branch coverage bitmap using the global seed corpus, and initialize the branch coverage bitmap of each fuzzer instance with this global branch coverage bitmap; ② During the fuzz testing process, read the branch coverage bitmap of each fuzzer instance at a fixed interval, add the newly covered branches into the global branch coverage bitmap, and then synchronize the global branch coverage bitmap to each fuzzer instance.

10. An integrated fuzz testing system based on function call graph partitioning, characterized in that It includes a global manager and fuzzers: The global manager is above each parallel heterogeneous fuzzer, and realizes the functions of task division, seed mapping and task assignment. The global manager contains five major modules: function call graph division module, seed mapping module, task management module, new seed division module and branch coverage bitmap synchronization module; The function call graph division module updates the function call graph of the target program extracted in advance using the execution paths of all seeds generated in the previous stage, simplifies it, and then divides it into subgraphs representing each subtask using the graph division algorithm; The seed mapping module divides the global seed corpus into sub-seed corpora corresponding one-to-one with the subtasks according to the association degree between the execution path of each seed in the global seed corpus and each subgraph; The task management module assigns subtasks to each parallel fuzzer instance, and rotates the fuzzer used by each subtask at a fixed interval; The new seed division module monitors the new seeds generated by each fuzzer instance in real time, and assigns them to the corresponding subtasks according to their execution paths; The branch coverage bitmap synchronization module maintains a global branch coverage bitmap, and synchronizes it with the branch coverage bitmap of each fuzzer instance itself regularly during the fuzz testing process; Corresponding to the new seed division module and the branch coverage bitmap synchronization module in the global management module, a seed synchronization policy and a branch coverage bitmap synchronization policy are added at the fuzzer side respectively, jointly ensuring that the fuzzer always keeps testing the part of the subtask assigned to it.

Citation Information

Patent Citations

  • Parallel fuzzy test method and system based on target point task division

    CN114328213A

  • Integrated fuzz testing method and system for automatically selecting fuzzer combination

    CN116010281A