A Dynamic Programming-based AFL Multithreaded Optimization Method and System
By adopting a multi-threaded optimization method based on dynamic programming in AFL, redundant paths in multi-threaded scheduling are removed, the seed stability of AFL in multi-threaded applications without target source code is improved, the problem of AFL's low efficiency in multi-threaded scheduling is solved, and a new solution is provided for intelligent vulnerability mining.
Patent Information
- Application Number
- CN202210740430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-06-27
AI Technical Summary
When the target program source code cannot be obtained, how to improve the stability of AFL seeds in multi-threaded scheduling, so that a test case corresponds to a path, solving the scheduling problem of multi-threaded applications without target source code.
AFL multi-threaded optimization method based on dynamic programming is adopted, and the test cases input by users is preprocessed and analyzed, and a new test case is generated using the AFL mutation algorithm, and the longest common subsequence node set of multiple paths is solved through the dynamic programming algorithm, redundant paths are removed, effective paths are obtained, and the stability of the seed queue is improved.
It significantly improves the stability of AFL seeds in multi-threaded scheduling, effectively solves the problem of AFL in multi-threaded applications without target source code, and provides a new solution for the application of intelligent vulnerability mining in multi-threaded scheduling.
Smart Images

Figure CN115186266B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and particularly relates to a dynamic programming-based AFL multi-thread optimization method and system. Background Art
[0002] Grey-box fuzz testing technology has played an important role in the field of vulnerability mining due to its advantages of lightweight detection, fast coverage feedback, and dynamic adjustment of strategies. One of the most effective technologies is coverage-guided fuzz testing represented by the fuzz testing tool AFL (American fuzzy lop). By using compile-time instrumentation, the coverage information during program execution can be obtained, which can improve the branch coverage of the code under test and help discover new vulnerabilities.
[0003] Although AFL has the above advantages, it usually faces single-thread problems, that is, assuming that the seeds are stable, a test case only covers one path, and its direct application to multi-thread problems is less efficient. However, in actual engineering applications, there are scenarios that require multi-threaded concurrent scheduling. The concurrent errors caused by multi-thread interleaving have a negative impact on the stability of the fuzz testing tool. Therefore, how to improve the seed stability of AFL when dealing with multi-threaded target programs is a problem that needs to be considered. Currently, some studies have proposed a solution to modify the program source code. By determining the execution order of program instructions, the thread concurrency is limited to the minimum as much as possible. However, this method only applies to the case where the source code of the target program is known. In engineering requirements, in most cases, the source code of the target program cannot be obtained. Therefore, improving the seed stability of AFL in multi-thread problems without the target program source code has become the main research goal. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] The technical problem to be solved by the present invention is: in the case where the source code of the target program cannot be obtained, how to improve the seed stability of AFL in multi-thread scheduling, so that a test case corresponds to one path, thereby solving the scheduling problem of multi-threaded application programs without the target source code.
[0006] (2) Technical Solutions
[0007] To solve the above technical problems, the present invention provides a dynamic programming-based AFL multi-thread optimization method, including the following steps:
[0008] S1. Preprocess the test cases input by the user and add the preprocessed test cases to the seed queue;
[0009] S2. Based on the seed queue, use the AFL mutation algorithm to obtain new test cases;
[0010] S3. Monitor whether the new test cases trigger new paths. If new paths are triggered, continue to step S4; otherwise, return to step S2.
[0011] S4. Determine whether the generated new paths are uniquely valid. If they are uniquely valid, transfer to step S6; if there are multiple new paths, continue to step S5.
[0012] S5. Eliminate redundant paths to obtain valid paths.
[0013] S6. Obtain the new test cases and their corresponding valid paths, add them to the seed queue, and obtain a valid set of test cases.
[0014] S7. Test the test cases and the selected valid paths, and then return to step S2 for the next round of new mutations until the test time ends or is manually stopped.
[0015] Preferably, step S1 is specifically: Analyze and test the test cases input by the user, judge the effectiveness and stability of the test cases until it is ensured that the test cases input by the user can trigger valid new paths, and then add the test cases to the seed queue.
[0016] Preferably, in step S1, the method for judging the effectiveness of the test cases is: Obtain the coverage path feedback of the program under test through compile-time instrumentation to judge whether the test cases are valid.
[0017] Preferably, step S2 is specifically: Take out the seeds with high priority from the seed queue, and use the AFL mutation algorithm to mutate the seeds to obtain new test cases generated by the seeds.
[0018] Preferably, step S3 is specifically: Test the new test cases generated by using the AFL mutation algorithm on the seeds, monitor whether new paths are triggered. If new paths are triggered, it means that the new test cases may be valid, and continue to step S4; otherwise, return to step S2 to continue generating the next new test case.
[0019] Preferably, step S4 is specifically: When the new path triggered by a test case is unique and can effectively discover the internal execution state of the program, it means that the test case is valid and stable and can be added to the seed queue, and then transfer to step S6. If multiple new paths are triggered, it is necessary to further judge whether there are redundant paths caused by multi-threaded interleaving, so transfer to step S5.
[0020] Preferably, step S5 is specifically:
[0021] Solve the longest common subsequence node set of all new path nodes based on dynamic programming algorithm;
[0022] Take the complement of all new path node sets and the longest common subsequence node set to obtain the interference node set, that is, take the path nodes that are not in the longest common subsequence to form the interference node set;
[0023] Remove interference nodes from all new paths, thereby eliminating redundant paths and obtaining effective coverage paths, namely, valid paths.
[0024] Preferably, in step S5, the longest common subsequence node set of all new path nodes is solved based on the dynamic programming algorithm, and the specific operation is:
[0025] Use a string to represent the set of nodes in each path, then N paths can be represented as X1, X2, ..., X N , Taking the longest common subsequence of paths X1 and X2 as an example, let C[i,j] be the longest common subsequence of the first i,j characters in paths X1 and X2, then the solution formula is:
[0026]
[0027] X1[i] and X2[j] represent the i-th and j-th characters in paths X1 and X2, respectively;
[0028] For multiple paths, the divide-and-conquer method is used to obtain them. All paths are divided into a fixed number of parts, and each part is further divided. After multiple backtracking, the longest common subsequence node set S of all paths is finally obtained.
[0029] Preferably, in step S5, the path nodes not in the longest common subsequence are selected to form an interference node set. The specific solution method is: from the path node set X1, X2, ..., X N The nodes in the longest common subsequence node set S are removed, and the final node set obtained is the interference node set.
[0030] The present invention also provides an AFL multi-thread optimization system based on dynamic programming for implementing the method, comprising a preprocessing module, a test case generation module, a path monitoring module, a path selection module, and a testing module;
[0031] The preprocessing module is used to preprocess the initial test case provided by the user, obtain the coverage path feedback of the tested program by means of compile-time instrumentation, so as to determine whether the test case is valid. If it is valid, the test case is added to the seed queue. If it is invalid, the user provides input again until a valid test case is obtained;
[0032] The test case generation module is used to obtain new test cases by relying on seed mutation. Specifically, seeds with high priority are first selected from the seed queue as the input of AFL, and then new test cases are generated by mutating the selected seeds according to the genetic algorithm.
[0033] The path monitoring module is used to monitor whether the new test cases can trigger new paths. If no new path is triggered, it returns to the test case generation module to continue generating the next test case. If a new path is triggered, all the triggered new paths are recorded and saved in detail.
[0034] The path selection module is used to remove redundant paths from the obtained paths to obtain valid paths when there are multiple new paths. Specifically, the longest common subsequence nodes of all paths are first solved based on the dynamic programming algorithm, and then the path nodes not in the longest common subsequence are used to form an interference node set. Next, the interference nodes in the paths are removed, and the redundant paths can be excluded. The remaining paths are selected as valid paths, and then the test case and its corresponding valid path are added to the seed queue.
[0035] The test module is used to test the test cases and the selected valid paths.
[0036] (3) Beneficial effects
[0037] The AFL multi-thread optimization method and system based on dynamic programming provided by the present invention solve the longest common subsequence node set of multiple path nodes based on the dynamic programming algorithm, and then find out the interference nodes. Using this method can effectively exclude redundant paths, significantly improve the stability of seeds in AFL multi-thread scheduling, and is beneficial to solving the low efficiency problem of AFL applications in multi-threaded applications without target source code, providing a new solution and research idea for the application problem of intelligent vulnerability mining in multi-thread scheduling. Description of the drawings
[0038] Figure 1 It is a flowchart of the AFL multi-thread optimization method and system based on dynamic programming provided by the present invention. Specific implementation manners
[0039] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.
[0040] The present invention provides a dynamic programming-based AFL multi-thread optimization method and system. This method and system target the multi-thread problem of AFL, with the dynamic programming algorithm as the basic guiding idea, aiming to remove the interference nodes in the coverage path, thereby obtaining an effective coverage path and improving the stability of seeds in AFL. The basic idea of this method and system is as follows: First, a set of effective test cases are provided by the user and added to the seed queue. Then, the AFL mutation algorithm is used to mutate the seeds to generate new test cases. For the redundant paths caused by the interleaving of multi-thread scheduling, the dynamic programming algorithm is used to solve the longest common subsequence node set of all paths. By obtaining the complement set of the combination of all path node sets and the longest common subsequence nodes, the interference nodes can be obtained, thereby eliminating the redundant paths and screening out the effective paths. Through step-by-step improvement and screening, an effective seed queue can finally be obtained, which is beneficial to improving the seed stability of AFL in multi-thread scheduling without the source code of the target program.
[0041] Reference Figure 1 , the dynamic programming-based AFL multi-thread optimization method of the present invention includes the following steps:
[0042] S1. Preprocess the test cases input by the user and add the preprocessed test cases to the seed queue;
[0043] S2. Based on the seed queue, use the AFL mutation algorithm to obtain new test cases;
[0044] S3. Monitor whether the new test cases trigger a new path. If a new path is triggered, continue to step S4; if not, return to step S2;
[0045] S4. Judge whether the generated new path is uniquely valid. If it is uniquely valid, transfer to step S6; if there are multiple new paths, continue to step S5;
[0046] S5. Eliminate redundant paths and obtain effective paths;
[0047] S6. Obtain the new test cases and their corresponding effective paths and add them to the seed queue;
[0048] S7. Test the test cases and the selected effective paths (if it can cause the program to crash, a potential vulnerability may be discovered, indicating the effectiveness of steps S1 to S6 of this optimization method), and then return to step S2 for the next round of new mutations until the test time ends or is manually stopped.
[0049] Step S1 is specifically:
[0050] Analyze and test the test cases input by the user to determine the effectiveness and stability of the test cases until it is ensured that the test cases input by the user can trigger valid new paths, and then add the test cases to the seed queue.
[0051] Among them, the method for determining the effectiveness of the test case is: obtain the coverage path feedback of the program under test through compile-time instrumentation to determine whether the test case is effective.
[0052] Step S2 is specifically as follows:
[0053] Take out the seeds with high priority from the seed queue, and mutate the seeds using a variety of AFL mutation algorithms to obtain new test cases generated from the seeds; the AFL mutation algorithm adopted in this embodiment is a genetic algorithm.
[0054] Step S3 is specifically as follows:
[0055] Test the new test cases generated from the seeds using the AFL mutation algorithm, monitor whether new paths are triggered. If new paths are triggered, it means that the new test cases may be effective, and then continue with step S4. If not, return to step S2 to continue generating the next new test case;
[0056] Step S4 is specifically as follows:
[0057] When the new path triggered by a test case is unique and can effectively discover the internal execution state of the program, it means that the test case is effective and stable and can be added to the seed queue, and then transfer to step S6. If multiple new paths are triggered, it is necessary to further determine whether there are redundant paths caused by multi-threaded interleaving, so transfer to step S5;
[0058] Step S5 is specifically as follows:
[0059] Solve the set of longest common subsequence nodes of all new path nodes based on the dynamic programming algorithm;
[0060] Take the complement of the set of all new path nodes and the set of longest common subsequence nodes to obtain the set of interfering nodes, that is, take the path nodes that are not in the longest common subsequence to form the set of interfering nodes;
[0061] Remove the interfering nodes from all new paths to exclude redundant paths and obtain the effective coverage path, that is, the effective path.
[0062] Among them, the specific operation of solving the set of longest common subsequence nodes of all new path nodes based on the dynamic programming algorithm is:
[0063] Represent the set composed of each path node as a string, then N paths can be represented as X1, X2,..., XN , taking the example of finding the longest common subsequence of paths X1 and X2, let C[i,j] be the longest common subsequence of the first i and j characters in paths X1 and X2, then its solution formula is:
[0064]
[0065] X1[i] and X2[j] represent the i-th and j-th characters in paths X1 and X2 respectively;
[0066] For multiple paths, the divide-and-conquer method is used to obtain them. All paths are divided into fixed parts, and each part is further divided. After multiple backtracks, the set S of the longest common subsequence nodes of all paths is finally obtained.
[0067] Among them, the path nodes that are not in the longest common subsequence are taken to form the interference node set. The specific solution method is: remove the nodes in the longest common subsequence node set S from the path node sets X1, X2,..., X N The final obtained node set is the interference node set.
[0068] Step S6 is specifically as follows:
[0069] After the screening of the above steps, effective test cases and their corresponding effective paths can be obtained. After adding them to the seed queue, an effective test case set is obtained.
[0070] The present invention also provides an AFL multi-thread optimization system based on dynamic programming, including a preprocessing module, a test case generation module, a path monitoring module, a path selection module, and a test module. The functions of each module are described in detail below.
[0071] (1) Preprocessing module
[0072] The preprocessing module is used to preprocess the initial test cases provided by the user, obtain the coverage path feedback of the program under test through compile-time instrumentation, and determine whether the test case is effective. If it is effective, the test case is added to the seed queue. If it is invalid, the user is required to provide input until an effective test case is obtained.
[0073] (2) Test case generation module
[0074] The test case generation module mainly obtains new test cases by seed mutation. Specifically, first, a seed with a high priority is selected from the seed queue as the input of AFL, and then the seed is mutated according to the genetic algorithm to generate new test cases.
[0075] (3) Path monitoring module
[0076] The path monitoring module is used to monitor whether a new test case can trigger a new path. If no new path is triggered, it returns to the test case generation module to continue generating the next test case. If a new path is triggered, all the triggered new paths are detailedly recorded and saved.
[0077] (4) Path selection module
[0078] The path selection module is used to remove redundant paths from the obtained paths to obtain valid paths when there are multiple new paths. Specifically, first, the longest common subsequence nodes of all paths are solved based on the dynamic programming algorithm. Then, the path nodes not in the longest common subsequence are taken to form a set of interference nodes. Next, the interference nodes in the paths are removed, so that the redundant paths can be excluded, and the remaining paths are selected as valid paths. Then, the test case and its corresponding valid path are added to the seed queue.
[0079] (5) Test module
[0080] The test module is used to test the test case and the selected valid path. If it can cause the program to crash, a potential vulnerability may be discovered, indicating the effectiveness of the optimization method.
[0081] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A method for optimizing AFL multi-threading based on dynamic programming, characterized in that, It includes the following steps: S1. Preprocess the test cases input by the user and add the preprocessed test cases to the seed queue; S2. Based on the seed queue, use the AFL mutation algorithm to obtain new test cases; S3. Monitor whether the new test cases trigger new paths. If new paths are triggered, continue with step S4. If not, return to step S2; S4. Judge whether the generated new paths are uniquely valid. If they are uniquely valid, proceed to step S6. If there are multiple new paths, continue with step S5; S5. Eliminate redundant paths to obtain valid paths; S6. Obtain new test cases and their corresponding valid paths, add them to the seed queue, and obtain a set of valid test cases; S7. Test the test cases and the selected valid paths, and then return to step S2 for the next round of new mutations until the test time ends or is manually stopped; Step S5 is specifically as follows: Solve the set of longest common subsequence nodes of all new path nodes based on the dynamic programming algorithm; Take the complement of the set of all new path nodes and the set of longest common subsequence nodes to obtain the set of interfering nodes, that is, take the path nodes not in the longest common subsequence to form the set of interfering nodes; Remove the interfering nodes from all new paths, thereby eliminating redundant paths and obtaining effective coverage paths, that is, valid paths; In step S5, the specific operation of solving the set of longest common subsequence nodes of all new path nodes based on the dynamic programming algorithm is as follows: Represent the set composed of each path node with a string, then the N paths can be represented as X1, X2, …, X N , taking the example of finding the longest common subsequence of paths X1 and X2, let C[i, j] be the longest common subsequence of the first i and j characters in paths X1 and X2, then its solution formula is: X1[i] and X2[j] respectively represent the i-th and j-th characters in paths X1 and X2; For multiple paths, the divide-and-conquer method is used for calculation. All paths are divided into fixed parts, and each part is further divided. After multiple backtracks, the set S of the longest common subsequence nodes of all paths is finally obtained.
2. The method according to claim 1, characterized in that, Step S1 is specifically as follows: Analyze and test the test cases input by the user, judge the effectiveness and stability of the test cases until it is ensured that the test cases input by the user can trigger valid new paths, and then add the test cases to the seed queue.
3. The method according to claim 2, characterized in that, In step S1, the method of judging the effectiveness of the test cases is: Obtain the coverage path feedback of the program under test through compile-time instrumentation to judge whether the test cases are valid.
4. The method according to claim 2, characterized in that, Step S2 is specifically as follows: Take out the seeds with high priority from the seed queue, and use the AFL mutation algorithm to mutate the seeds to obtain new test cases generated by the seeds.
5. The method according to claim 4, characterized in that, Step S3 is specifically as follows: Test the new test cases generated by using the AFL mutation algorithm on the seeds, monitor whether new paths are triggered. If new paths are triggered, it means that the new test cases may be valid, and continue with step S4. If not, return to step S2 to continue generating the next new test case.
6. The method according to claim 5, characterized in that, Step S4 is specifically as follows: When the new path triggered by a test case is unique and can effectively discover the internal execution state of the program, it indicates that the test case is effective and stable and can be added to the seed queue, then go to step S6. If multiple new paths are triggered, it is necessary to further determine whether there are redundant paths caused by multi-thread interleaving, so go to step S5.
7. The method according to claim 1, characterized in that, In step S5, the path nodes that are not in the longest common subsequence are taken to form a set of interference nodes. The specific solution method is as follows: Remove the nodes in the longest common subsequence node set S from the path node sets X1, X2, …, X N to finally obtain the node set as the set of interference nodes.
8. A system for optimizing AFL multi-threading based on dynamic programming for implementing the method according to claim 7, characterized in that, It includes a preprocessing module, a test case generation module, a path monitoring module, a path selection module, and a test module; The preprocessing module is used to preprocess the initial test case provided by the user, obtain the coverage path feedback of the program under test through compile-time instrumentation, and determine whether the test case is effective. If it is effective, the test case is added to the seed queue. If it is not effective, the user is required to provide input until an effective test case is obtained; The test case generation module is used to obtain new test cases by relying on seed mutation. Specifically, first, select the seed with the highest priority from the seed queue as the input of AFL, and then mutate the seed according to the genetic algorithm to generate new test cases; The path monitoring module is used to monitor whether a new test case can trigger a new path. If no new path is triggered, return to the test case generation module to continue generating the next test case. If a new path is triggered, record and save all the triggered new paths; The path selection module is used to remove redundant paths from the obtained paths to obtain effective paths when there are multiple new paths. Specifically, first, solve the longest common subsequence nodes of all paths based on the dynamic programming algorithm, then select the path nodes that are not in the longest common subsequence to form an interference node set, and then remove the interference nodes from the paths to eliminate the redundant paths. Select the remaining paths as effective paths, and then add the test case and its corresponding effective path to the seed queue; The test module is used to test the test case and the selected effective path.
Citation Information
Patent Citations
Generative adversarial network-based AFL seed optimization method and system
CN115391787A