High code coverage rate fuzzy testing method based on parallel strategy
By deploying multiple parallel instances in fuzz testing, combining dynamic instrumentation and central management system, optimizing file operations and multi-threaded synchronization, the shortcomings of existing fuzz testing methods in parallel strategies are solved, and efficient code coverage and vulnerability discovery are achieved.
Patent Information
- Application Number
- CN202510398209.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
AI Technical Summary
The existing fuzz testing methods have shortcomings in efficiency and coverage, especially in the lack of effective optimization solutions in parallel strategies, resulting in low resource utilization and low testing efficiency.
The high-code coverage fuzz testing method based on parallel strategies is adopted. By deploying multiple parallel running fuzz testing instances, combining dynamic instrumentation technology to collect the execution path and status data of the target program in real time, dynamically evaluate seed priority based on the central management system, and allocate high-priority seeds through the feedback driver mechanism, optimize file operations and multi-thread synchronization, and dynamically adjust seed priority and test strategy.
It significantly improves the efficiency and coverage of fuzz testing, reduces system calls and resource competition, realizes efficient collaboration and intelligent testing strategy adjustments, and improves code coverage and vulnerability discovery capabilities.
Smart Images

Figure CN120295922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fuzz testing, and more precisely, it relates to a high code coverage fuzz testing method based on a parallel strategy. Background Art
[0002] As software systems become increasingly complex, single-threaded or traditional fuzz testing methods face challenges in terms of efficiency and coverage. The technology of fuzz testing has started to develop in more dimensions. For example, random fuzz testing has begun to be replaced by structured fuzz testing, which can generate test cases based on the specific implementation and logic of the software. Godefroid et al. introduced a technology called "white-box fuzz testing", which is actually a structured fuzz testing method. It uses a combination of static analysis and dynamic testing to automatically generate test cases, focusing on program logic and implementation structure. This method can more effectively cover code paths, thereby increasing the probability of discovering vulnerabilities. In addition, the EXE tool proposed by Cadar generates test cases through symbolic execution, which can accurately cover different execution paths of the program, thus effectively discovering potential vulnerabilities.
[0003] Furthermore, feedback-based fuzz testing methods have also emerged. They dynamically adjust test inputs by monitoring the running state of the program in real time to improve the effectiveness of testing. AFL is a widely used feedback-based fuzz testing tool. It monitors the running state of the program and dynamically adjusts test inputs to improve the effectiveness and efficiency of testing. AFL can discover new paths in program execution and generate new test cases based on this, further increasing coverage. The desire to increase code coverage is one of the core goals in improving fuzz testing efficiency. To achieve this core goal, researchers often adopt two main types of optimization strategies. The first type of strategy is to improve the algorithmic process of fuzz testing technology, modifying the deficiencies of the algorithm and software characteristics in order to improve testing efficiency. The second type of strategy is to use parallel testing, by starting multiple fuzz testing instances in parallel to increase the number of tests per unit time, thereby increasing code coverage. With the development of computing technology, various scholars have proposed many optimization algorithms in the direction of the first type of strategy and achieved remarkable progress. However, there is currently a lack of sufficient discussion in the direction of the second type of strategy. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies of the prior art and propose a high code coverage fuzz testing method based on a parallel strategy.
[0005] In a first aspect, there is provided a high code coverage fuzz testing method based on a parallel strategy, including:
[0006] S1. Deploy multiple fuzz testing instances running in parallel;
[0007] S2. Real-time collect the execution path and status data of the target program through dynamic instrumentation technology;
[0008] S3. Dynamically evaluate the seed priority based on the central management system, and the priority is determined according to the uniqueness of the execution path, the new path discovery rate, and the historical vulnerability trigger effectiveness;
[0009] S4. Allocate high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimize file operations and multithreaded synchronization;
[0010] S5. Dynamically adjust the seed priority and testing strategy according to the real-time test data.
[0011] Preferably, in S2, the dynamic instrumentation technology is implemented through LLVM IR Pass, including: inserting instructions using the IRBuilder tool of LLVM, recording the basic block jump information and path identifiers of the program, and feeding the data back to the central management system in real time.
[0012] Preferably, in S3, the dimension indicators for evaluating the seed priority include: the trigger path of the seed, the number of new paths, and the historical performance.
[0013] Preferably, in S5, the dynamic adjustment of the seed priority includes:
[0014] When a new path or vulnerability is found in an instance, the central management system raises the priority of the relevant seeds and triggers global synchronization to accelerate the coverage of similar paths; seeds that fail to produce effective results for a long time are then de-prioritized or temporarily removed from the test queue.
[0015] Preferably, in S3, the central management system collects the running data from each fuzz testing instance, and the running data includes the execution path, error log, and performance metrics.
[0016] In a second aspect, a high code coverage fuzz testing system based on a parallel strategy is provided for executing any of the methods in the first aspect, including:
[0017] A deployment module for deploying multiple parallel-running fuzz testing instances;
[0018] A collection module for real-time collecting the execution path and status data of the target program through dynamic instrumentation technology;
[0019] An evaluation module for dynamically evaluating the seed priority based on the central management system, and the priority is determined according to the uniqueness of the execution path, the new path discovery rate, and the historical vulnerability trigger effectiveness;
[0020] An allocation module for allocating high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimizing file operations and multithreaded synchronization;
[0021] An adjustment module for dynamically adjusting the seed priority and testing strategy according to real-time test data.
[0022] In a third aspect, a computer storage medium is provided, in which a computer program is stored; when the computer program runs on a computer, the computer is caused to execute the method according to any one of the first aspects.
[0023] In a fourth aspect, an electronic device is provided, including:
[0024] A memory for storing a computer program;
[0025] A processor for executing the computer program to implement the method according to any one of the first aspects.
[0026] The beneficial effects of the present invention are as follows: By optimizing the parallel seed scheduling algorithm and file operation process, and combining with the LLVM IR Pass instrumentation technology, the present invention not only effectively reduces system calls and resource contention, improves test efficiency and coverage rate, but also realizes efficient cooperation and intelligent adjustment of testing strategies through the multithreaded synchronization mechanism and the design of the central management system, overcomes the deficiencies of the prior art in terms of performance, complexity and resource utilization, and demonstrates more excellent performance and practicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flowchart of a high code coverage fuzz testing method based on a parallel strategy provided by the present application;
[0028] Figure 2 It is a schematic diagram of the parallel strategy provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described below with reference to embodiments. The description of the following embodiments is only for helping to understand the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0030] Embodiment 1:
[0031] AFL is a fuzz testing tool widely used in the industry and has successfully discovered multiple vulnerabilities in production environments. In its parallel mode, the system starts multiple AFL instances simultaneously, and each instance maintains its own test information in a common directory for storing the test cases generated during the test. The designers hope that the parallel performance of AFL will increase linearly with the increase of resources. However, if each instance runs independently without interaction, the performance improvement will be very limited. Therefore, AFL introduces a key step of guiding information synchronization to enable each instance to
[0032] obtain valuable test case files to facilitate more effective collaboration in completing the test tasks. This synchronization mechanism has become the basis for subsequent research on parallel fuzz testing, and the synchronization accuracy is also continuously improving, including code coverage and crash information, etc.
[0033] AFL's synchronization mechanism allows each instance to scan the local directories of other instances after completing the regular test loop to obtain valid test cases. To avoid repeated scanning of the same content, AFL assigns a strict serial number to each test case and records the serial number of the last test case scanned each time as the starting point for the next scan, which is called the synchronization phase. During the scan, the instance will execute the test cases newly generated by other instances in sequence and collect the execution feedback. If the executed test case touches a new logical branch, the instance will save the test case in the local directory for further use. This synchronization design allows different instances to share valuable test cases and reduces redundant exploration in the huge input space. However, the computational complexity of each instance traversing the file directories of other instances in the algorithm reaches O(n*m), where n is the number of fuzz testing instances and m is the number of test cases in the test case directory. This seriously affects the parallel performance. In addition, AFL completely does not consider the optimization scheme for seed synchronization information of multiple instances, and thus cannot perform relevant algorithm optimization on the repetition rate of different seeds.
[0034] To solve the problems of the existing technology, this application provides a high code coverage fuzz testing method based on a parallel strategy. The present invention mainly explores parallel fuzz testing technology and is committed to significantly improving the efficiency and coverage of fuzz testing through parallel computing technology. This technology is particularly suitable for application in large-scale and complex basic software programs. The core research content of the present invention covers the development of an efficient fuzz testing algorithm, the design and implementation of a system prototype, aiming to improve the efficiency of parallel fuzz testing and discover program application vulnerabilities faster.
[0035] Specifically, as Figure 1 shown, the high code coverage fuzz testing method based on a parallel strategy includes:
[0036] S1. Deploy multiple fuzz testing instances running in parallel.
[0037] The core idea of the parallel fuzzing strategy is to run multiple fuzzing instances simultaneously to improve the efficiency and coverage of the testing process. This method is particularly effective when dealing with large software projects because it can significantly speed up the vulnerability identification. For example, by running multiple test instances, researchers can explore different parts of the software simultaneously, which is difficult to achieve in single-instance fuzzing.
[0038] S2. Real-time collect the execution path and status data of the target program through dynamic instrumentation technology.
[0039] In S2, the dynamic instrumentation technology is implemented through LLVM IR Pass, including: using the IRBuilder tool of LLVM to insert instructions, record the basic block jump information and path identifiers of the program, and feedback the data to the central management system in real time.
[0040] This application optimizes the core seed writing function of AFL, simplifies the file operation process, and reduces the file system calls, which is crucial for improving the parallel processing speed. The prior art completely fails to consider the optimization scheme for the seed synchronization information of multiple instances, and thus cannot optimize the relevant algorithms for the duplication rate of different seeds. Therefore, it is very important to use llvm IRPass to re-traverse and obtain relevant information.
[0041] Specifically, in the core read-write function write_to_testcase of AFL, the main function of the code is to write data to a specified test case file. This includes operations such as file opening, writing, truncating, and closing. Although this code is functionally complete, there are some potential efficiency and error handling issues. First, the simultaneous presence of unlink and O_EXCL may lead to unnecessary system calls. Second, attempting to delete and recreate the file every time data is written may cause performance degradation during high-frequency writing. The optimized code has made some key improvements in processing logic and performance, removing the unlink call and O_EXCL flag. This reduces the complexity of file system operations and the number of system calls, and also avoids additional error handling due to the existence of the file. By reducing unnecessary file deletion and reconstruction operations, the optimized code may improve the performance of file operations, especially in the case of frequent file writing.
[0042] In addition, the stub code of the present invention uses the IRBuilder Pass tool of LLVM to construct instructions. These instructions include loading the identifier of the previous position, performing necessary type conversions, and calling the printing function to output the current and previous position information. This step is crucial because it helps us track the state changes during program execution. To maintain the continuity of the execution state, we also update the global variable storing the identifier of the previous position by shifting the current position to the right. Such a design not only helps us monitor the dynamic behavior of the program but also promotes the understanding of the program execution path, especially in dealing with complex code bases and multi-threaded application scenarios. Through such dynamic code insertion and execution monitoring strategies, our system can capture the execution details of the program in real time, which is crucial for identifying and debugging potential execution problems, thus enhancing the effectiveness of fuzzing tests.
[0043] S3. Dynamically evaluate the seed priorities based on the central management system, where the priorities are determined according to the uniqueness of the execution path, the new path discovery rate, and the effectiveness of historical vulnerability triggers.
[0044] In S3, the dimensional metrics for evaluating the seed priorities include: the trigger path of the seed, the number of new paths, and the historical performance.
[0045] Since the states and attributes of seeds may change frequently, a system that can respond quickly to these changes is needed. The role of the priority queue here is to ensure that seeds with high priorities (i.e., those with high path complexity, good historical performance, or newly discovered paths) can be processed first, which can increase the depth and breadth of the test and is more likely to uncover undetected defects in the software. The hash table is used to quickly search for and manage the metadata of seeds, such as their unique identifiers, current states, code coverage, etc., which is essential for implementing complex scheduling logic and maintaining a large-scale seed library.
[0046] S4. Allocate high-priority seeds to each of the fuzzing test instances through a feedback-driven mechanism, and optimize file operations and multi-thread synchronization.
[0047] S5. Dynamically adjust the seed priorities and test strategies according to the real-time test data.
[0048] Embodiment 2:
[0049] Based on Embodiment 1, Embodiment 2 of the present application provides an efficient parallel seed scheduling algorithm. Based on the data structures of a priority queue and a hash table, combined with a multi-thread synchronization mechanism, it realizes efficient seed management and scheduling, significantly improving the coverage rate and efficiency of parallel fuzz testing. Secondly, in the core read and write functions of AFL, redundant file deletion and reconstruction operations are removed, reducing the number of system calls, improving the performance in high-frequency write scenarios, and enhancing the stability and response speed of the system. In addition, a dynamic instrumentation technique based on LLVM IR Pass is adopted to precisely instrument the target program, real-time tracking the program execution path and state changes, providing detailed data support for feedback-driven seed scheduling. Furthermore, a central management system is introduced, which is responsible for collecting and summarizing the running data of each fuzz testing instance, dynamically adjusting the priority and allocation strategy of seeds, and realizing efficient state synchronization and resource coordination. At the same time, through a feedback-driven dynamic scheduling mechanism, the execution results of fuzz testing instances are monitored in real time, and the selection and priority of seeds are dynamically adjusted to ensure that high-value seeds are processed preferentially, increasing the probability of vulnerability discovery and the depth of testing. In a multi-thread environment, a combination of fine-grained locks and lock-free data structures is adopted to ensure data consistency and system stability during high-concurrency access, further improving the execution efficiency of parallel fuzz testing. Finally, based on multi-dimensional metrics such as the trigger path of seeds, the number of new paths, and historical performance, a dynamic priority evaluation and adjustment strategy is designed to ensure that high-value seeds can be processed preferentially, optimizing the utilization rate of test resources. Through these innovative points, the present invention provides an efficient and intelligent solution in the field of parallel fuzz testing, significantly improving the test coverage rate and vulnerability discovery ability, and having important technical value and practical significance.
[0050] Specifically, the high-code-coverage fuzz testing method based on a parallel strategy provided by the present application includes:
[0051] S1. Deploy multiple fuzz testing instances running in parallel.
[0052] S2. Real-time collect the execution path and state data of the target program through a dynamic instrumentation technique.
[0053] S3. Dynamically evaluate the seed priority based on a central management system, and the priority is determined according to the uniqueness of the execution path, the new path discovery rate, and the historical vulnerability trigger effectiveness.
[0054] In S3, the central management system collects the running data from each fuzz testing instance, and the running data includes the execution path, error log, and performance metrics.
[0055] S4. Allocate high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimize file operations and multi-thread synchronization.
[0056] In multi-threaded implementations, the role of synchronization mechanisms cannot be ignored. Locks and other synchronization tools are used in the algorithm to ensure data consistency and the stable operation of the algorithm when accessing and modifying shared resources in a multi-threaded environment. For example, when updating the seed state or recalculating the priority of a seed, the atomicity of the operation must be ensured to avoid data races or state conflicts. Although synchronization mechanisms may introduce certain performance overheads, through well-designed lock strategies and the avoidance of unnecessary lock contention, this overhead can be effectively controlled. Another key component of the algorithm is the feedback loop, which dynamically adjusts the seed selection and allocation strategies based on the execution results of each fuzzing instance. This adaptive mechanism based on real-time feedback enables the algorithm to adjust its behavior in real time to cope with changing test environments and program behaviors.
[0057] S5. Dynamically adjust the seed priority and test strategy according to real-time test data.
[0058] For example, if a certain instance discovers a new important path or vulnerability, the priority of the corresponding seed and its related seeds can be immediately increased, while those seeds that have not produced effective results for a long time may have their priorities decreased or be temporarily removed from the test queue. Such a strategy ensures the reasonable allocation and utilization of resources, accelerates the process of important discoveries. Ultimately, the combination of these technologies and strategies makes the parallel seed scheduling algorithm not only improve the efficiency of fuzzing testing but also enhance the ability to discover software defects, which has a significant effect on improving the quality and security of software. This strategy, through precise control and resource allocation, ensures that test resources are not wasted on unnecessary tests but are concentrated on those test cases that are most likely to reveal new errors.
[0059] In terms of fuzzing test state synchronization, this application designs a central management system that is responsible for collecting the running data from each fuzzing instance, including execution paths, error logs, and performance metrics. These data are used to evaluate the effectiveness of each seed and determine its future direction, such as whether to continue using the seed for in-depth testing or to adjust its priority in the queue. Generally speaking, this parallel seed scheduling algorithm is implemented through a series of well-designed strategies and technologies, effectively improving the efficiency of fuzzing testing and the ability to discover software defects. It not only balances the load, optimizes resource utilization, but also improves the intelligence and adaptability of the testing process, which has important practical application value for modern software development and testing.
[0060] It should be noted that for the technical goal of the present invention aiming to improve the efficiency and coverage of parallel fuzz testing, although the present invention provides an efficient parallel seed scheduling algorithm and its optimization measures, there are still other feasible alternative solutions. These solutions, while achieving similar goals, also have certain limitations. For example, using a distributed computing framework to distribute fuzz testing tasks to multiple physical or virtual nodes for execution can indeed utilize distributed resources to improve the parallelism and scalability of testing. However, it introduces higher network communication overhead and system complexity, and requires complex coordination mechanisms to ensure data consistency and the efficiency of task scheduling. In addition, although the seed selection strategy based on machine learning can predict the potential of high-value seeds by analyzing the data generated during the testing process and optimize the selection and scheduling of seeds, it relies on a large amount of training data and computing resources, and the accuracy of the model directly affects the testing effect, resulting in problems such as high implementation difficulty and cost. Another alternative is to adopt a hybrid fuzz testing method that combines symbolic execution with fuzz testing. By using the precision of symbolic execution to generate high-quality seeds, it makes up for the lack of randomness in traditional fuzz testing. However, the computational overhead of symbolic execution is relatively large, which limits its application to large-scale programs, and the efficiency of symbolic execution and its cooperation mechanism with fuzz testing need to be optimized.
[0061] Implementing the above technical solutions, the present invention has achieved a number of significant beneficial effects in the field of parallel fuzz testing. First of all, the optimized parallel seed scheduling algorithm and file operation process significantly reduce system calls and resource contention, thereby improving the overall execution speed of fuzz testing and significantly enhancing the testing efficiency. Secondly, through priority evaluation and feedback-driven dynamic scheduling, it ensures that high-value seeds can be processed preferentially, covering more code paths, improving the comprehensiveness and depth of testing, and thus increasing the code coverage. In addition, the design of the multi-thread synchronization mechanism and the central management system ensures that each fuzz testing instance can cooperate efficiently, quickly discover potential program vulnerabilities and security defects, and significantly enhance the vulnerability discovery ability. The refined seed management and scheduling strategy enables reasonable allocation of testing resources, avoids resource waste, and improves the resource utilization rate of parallel fuzz testing. The optimized error handling mechanism and synchronization mechanism enhance the stability and robustness of the system in a high-concurrency environment, ensuring the smooth progress of the fuzz testing process. At the same time, combining the LLVM IR Pass instrumentation technology and the real-time feedback mechanism of the central management system enables the system to dynamically adjust the testing strategy according to changes in the testing environment and program behavior, significantly improving the intelligence and adaptability of fuzz testing. These beneficial effects not only improve the efficiency and coverage of fuzz testing, but also enhance the stability and intelligence level of the system, providing a strong guarantee for the improvement of software quality and security.
[0062] It should be noted that the same or similar parts in this embodiment and Embodiment 1 can be referred to each other, and will not be elaborated in this application.
[0063] Embodiment 3:
[0064] Based on Embodiment 2, Embodiment 3 of this application provides a high code coverage fuzz testing system based on a parallel strategy, including:
[0065] A deployment module for deploying multiple fuzz testing instances running in parallel;
[0066] A collection module for real-time collecting the execution paths and status data of the target program through dynamic instrumentation technology;
[0067] An evaluation module for dynamically evaluating the seed priorities based on a central management system, where the priorities are determined according to the uniqueness of the execution paths, the new path discovery rate, and the effectiveness of historical vulnerability triggers;
[0068] An allocation module for allocating high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimizing file operations and multithreaded synchronization;
[0069] An adjustment module for dynamically adjusting the seed priorities and testing strategies according to real-time test data.
[0070] It should be noted that the system provided in this embodiment is the system corresponding to the method provided in Embodiment 2. Therefore, the same or similar parts in this embodiment and Embodiment 2 can be referred to each other, and will not be elaborated in this application.
Claims
1. A high code coverage fuzz testing method based on a parallel strategy, characterized in that Including: S1. Deploy multiple fuzz testing instances running in parallel; S2. Real-time collect the execution paths and status data of the target program through dynamic instrumentation technology; S3. Dynamically evaluate the seed priorities based on the central management system, and the priorities are determined according to the uniqueness of the execution path, the new path discovery rate, and the effectiveness of historical vulnerability triggering; S4. Allocate high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimize file operations and multithreaded synchronization; S5. Dynamically adjust the seed priorities and testing strategies according to the real-time test data.
2. The high-code-coverage fuzz testing method based on a parallel strategy according to claim 1, characterized in that In S2, the dynamic instrumentation technology is implemented through LLVM IR Pass, including: inserting instructions using the IRBuilder tool of LLVM, recording the basic block jump information and path identifiers of the program, and feeding the data back to the central management system in real time.
3. The high code coverage fuzz testing method based on a parallel strategy according to claim 2, wherein, In S3, the dimension indicators for evaluating the seed priorities include: the triggering paths of the seeds, the number of new paths, and the historical performance.
4. The high-code-coverage fuzz testing method based on a parallel strategy according to claim 3, wherein In S5, the dynamic adjustment of the seed priorities includes: When a new path or vulnerability is discovered by an instance, the central management system raises the priorities of the relevant seeds and triggers global synchronization to accelerate the coverage of similar paths; seeds that fail to produce effective results for a long time are then de-prioritized or temporarily removed from the test queue.
5. The high code coverage fuzz testing method based on a parallel strategy according to claim 4, wherein In S3, the central management system collects the running data from each fuzz testing instance, and the running data includes execution paths, error logs, and performance metrics.
6. A high code coverage fuzz testing system based on a parallel strategy, characterized in that, For executing the method according to any one of claims 1 to 5, including: A deployment module for deploying multiple fuzz testing instances running in parallel; A collection module for real-time collecting the execution paths and status data of the target program through dynamic instrumentation technology; An evaluation module for dynamically evaluating the seed priorities based on the central management system, and the priorities are determined according to the uniqueness of the execution path, the new path discovery rate, and the effectiveness of historical vulnerability triggering; An allocation module for allocating high-priority seeds to each of the fuzz testing instances through a feedback-driven mechanism, and optimizing file operations and multithreaded synchronization; An adjustment module for dynamically adjusting the seed priorities and testing strategies according to the real-time test data.
7. A computer storage medium, characterized in that, The computer storage medium stores a computer program; when the computer program runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 5.
8. An electronic device, characterized in that, Including: A memory for storing the computer program; A processor for executing the computer program to implement the method according to any one of claims 1 to 5.
Citation Information
Cited By
Vulnerability mining method and device for intelligent unmanned system controller program
CN121278725A