A vulnerability mining method and related device based on large models and fuzz testing
Through vulnerability mining methods based on large models and fuzzy testing, the problems of false alarms and inefficient vulnerability verification in the existing technology are solved, efficient and accurate vulnerability mining and verification are achieved, and the security of the software is ensured.
Patent Information
- Application Number
- CN202411532704.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-30
AI Technical Summary
The existing vulnerability mining methods have the problems of frequent false alarms, complex writing of vulnerability verification scripts, and inefficient fuzz testing.
Vulnerability mining methods based on large models and fuzzy testing are adopted, calling paths are obtained through static analysis, driver functions are generated using large models, and targeted fuzz testing is carried out in combination with vulnerability verification scripts and constraints to gradually mine vulnerabilities.
It improves the efficiency and accuracy of vulnerability mining, enhances the degree of automation, expands the coverage of vulnerabilities, and ensures software security.
Smart Images

Figure CN119397551B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to a vulnerability mining method, and particularly relates to a vulnerability mining method and related devices based on large models and fuzz testing. Background Art
[0002] With the continuous improvement and expansion of the software supply chain ecosystem, the continuous increase in the number of vulnerabilities in open-source components in the open-source software supply chain, and the increasing complexity of the dependencies between software components, the security issues caused by the open-source software supply chain have attracted more and more attention. During the software development process, due to the reasons of avoiding redundant development and improving software development efficiency, software developers are gradually introducing a large number of open-source third-party components into their own projects. However, these components are usually a "black box" to developers, and there may be thousands or tens of thousands of projects that depend on an open-source component. This leads to the situation that when a vulnerability is found in a component, developers of downstream projects cannot predict it in advance, and all software that depends on these components will face the same security threat. Further, the same vulnerability can be simply reproduced in different software, which may lead to large-scale economic losses. In addition, the affected software itself may also be a dependent component of other projects, resulting in the exponential spread of the resulting security issues.
[0003] Threats from the open-source software supply chain have prompted developers to adopt means such as software composition analysis to detect vulnerabilities in software, and use methods such as manual writing or fuzz testing to construct vulnerability verification scripts, so as to mine vulnerabilities and then prove the existence of vulnerabilities. However, these means have the following deficiencies: frequent false alarms, complex writing of vulnerability verification scripts, low efficiency of single fuzz testing, etc. Summary of the Invention
[0004] In view of the technical problems of frequent false alarms, complex writing of vulnerability verification scripts, and low efficiency of single fuzz testing existing in the existing vulnerability mining methods, this application provides a vulnerability mining method and related devices based on large models and fuzz testing.
[0005] To achieve the above object, this application is implemented by adopting the following technical solutions:
[0006] In the first aspect, this application proposes a vulnerability mining method based on large models and fuzz testing, including:
[0007] S1, performing static analysis on the target software, and combining the entry function information and vulnerability function information to obtain all call paths from the entry function to the vulnerability function in a forward direction;
[0008] S2. Calculate and sort the complexity of all call paths to obtain a set of call path queues; successively execute steps S3 to S7 on the call paths in the call path queue according to the ascending complexity of the call paths;
[0009] S3. Input the Prompt template and the call scenario of the call function in the call path into the large model to guide the generation of a driver function that can run the call function;
[0010] S4. If the driver function cannot run the call function normally, within the limited number of optimization times, feedback the running result to the large model to optimize the driver function through the large model;
[0011] S5. Combine the driver function, input the vulnerability verification script of the existing vulnerable function and the Prompt template information into the large model, use the large model to analyze the constraint conditions, and select the input parameters that need to be mutated;
[0012] S6. Generate seeds, and based on the input parameters that need to be mutated, conduct targeted fuzz testing on the call function to discover vulnerabilities and obtain the vulnerability verification script of the call function;
[0013] S7. Use the vulnerability verification script of the call function as a new reference, and recursively reverse-discover the vulnerabilities of the call functions on the current call path in a loop until the vulnerabilities of the entry function are discovered.
[0014] Further, the forward acquisition of all call paths from the entry function to the vulnerable function includes:
[0015] Use static analysis technology to analyze the target software to obtain a call graph. Based on the call graph, starting from the entry function and ending at the vulnerable function, use a graph traversal algorithm to obtain the set of all function nodes and call edges from the start point to the end point, and forwardly obtain all call paths.
[0016] Further, the calculation and sorting of the complexity of all call paths to obtain a set of call path queues includes:
[0017] Traverse each call path and generate the control flow graph of each call function. Starting from the entry function, successively count the number of nodes, the number of branch paths, the number of loop structures, the number of nested loops, and the maximum nested loop depth between the current call function and the next call function on the call path, and record them as inter-function parameters;
[0018] According to the control flow graph and the inter-function parameters, calculate the cyclomatic complexity between two call functions using the point-edge calculation method, and perform weighted summation together with the complexity generated by the loop structure to obtain the code complexity from the current function to the next adjacent call function;
[0019] Sum the complexities of all called functions and calculate the arithmetic mean;
[0020] Perform a weighted sum of the arithmetic mean and the number of functions in each call path to calculate the complexity of each call path.
[0021] Furthermore, the code complexity from the current function to the next immediately called function is obtained through the following formula:
[0022]
[0023] where, is the complexity of function , and are the cyclomatic complexity and the weight of the loop structure respectively, is the number of edges in the control flow graph, is the number of nodes in the control flow graph, , , are the number of loop structures, the number of nested loop structures, and the maximum nesting depth respectively.
[0024] Furthermore, the process of exploiting vulnerabilities to obtain the vulnerability verification script for the called function includes:
[0025] Set the parameter to be mutated as the input parameter to be mutated, and fix the parameters that remain unchanged;
[0026] Generate an initial seed set, and perform directed fuzz testing on the called function based on the initial seed set to obtain the vulnerability verification script for the called function.
[0027] Furthermore, the process of performing directed fuzz testing on the called function based on the initial seed set to obtain the vulnerability verification script for the called function includes:
[0028] During fuzz testing, calculate the distance between the nodes covered by the seed and the vulnerable function, and assign a higher priority to the seed with a shorter distance;
[0029] Continue to perform fuzz testing to exploit vulnerabilities until the end condition is met: exceeding the limited running time or finding an input that can verify the existence of vulnerabilities in the called function;
[0030] Combine the current input with the driver function to obtain the corresponding vulnerability verification script.
[0031] Furthermore, the process of recursively reverse-exploiting the vulnerabilities of the called functions on the current call path until the vulnerabilities of the entry function are exploited includes:
[0032] Take the vulnerability verification script of the calling function as a new reference, and regard the current calling function as a new vulnerable function;
[0033] Retrospectively select the next adjacent function in the calling path as the calling function, and repeat steps S3 to S6 to generate the corresponding vulnerability verification script;
[0034] Taking the entry function as the end point, continuously loop and recursively mine the vulnerability verification scripts of each calling function until the vulnerability of the entry function is mined.
[0035] In a second aspect, the present application proposes a vulnerability mining system based on a large model and fuzz testing, including:
[0036] A static analysis module for statically analyzing the target software, and forward obtaining all calling paths from the entry function to the vulnerable function by combining the entry function information and the vulnerable function information;
[0037] A mining module for calculating and sorting the complexity of all calling paths to obtain a set of calling path queues; and successively execute the following steps for the calling paths in the calling path queue from low to high complexity:
[0038] Input the Prompt template and the calling scenario of the calling function in the calling path into the large model to guide the generation of a driving function that can run the calling function;
[0039] If the driving function cannot run the calling function normally, within the limited number of optimization times, feedback the running result to the large model to optimize the driving function through the large model;
[0040] Combine the driving function, input the vulnerability verification script of the existing vulnerable function and the Prompt template information into the large model, and use the large model to analyze the constraint conditions to select the input parameters that need to be mutated;
[0041] Generate seeds, and based on the input parameters that need to be mutated, perform directed fuzz testing on the calling function to mine vulnerabilities and obtain the vulnerability verification script of the calling function;
[0042] Take the vulnerability verification script of the calling function as a new reference, and recursively mine the vulnerabilities of the calling functions on the current calling path in a loop until the vulnerabilities of the entry function are mined.
[0043] In a third aspect, the present application proposes an electronic device, including: a memory, one or more processors; the memory is coupled to the processors; wherein, the memory stores computer program code, and the computer program code includes computer instructions, when the computer instructions are executed by the processors, the electronic device executes the steps of the above-mentioned vulnerability mining method based on a large model and fuzz testing.
[0044] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned vulnerability mining method based on large models and fuzz testing.
[0045] Compared with the prior art, the present application has the following beneficial effects:
[0046] The present application proposes a vulnerability mining method based on large models and fuzz testing, introducing a method of calculating and sorting the complexity of call paths using static analysis techniques. By calculating and sorting the complexity of call paths and preferentially processing the paths with the lowest complexity, the overall process is significantly accelerated, and the efficiency of vulnerability mining is remarkably improved. When generating the driver function, the large model is combined with an automated repair mechanism. Even if the generated driver function fails to execute, it can reduce manual intervention through automatic feedback and iterative optimization, thereby enhancing the automation of the overall vulnerability mining process. By using the large model to analyze the constraints that need to be bypassed to trigger vulnerabilities before fuzz testing, combined with existing vulnerability verification scripts to select the parameters that need to be mutated in fuzz testing, and finally performing targeted fuzz testing with the vulnerable function as the target, it makes up for the deficiencies of large randomness and lack of purpose in fuzz testing, greatly improves the accuracy of fuzz testing, and increases the success rate of vulnerability mining. Finally, by using the successfully generated vulnerability verification script as a new reference to perform reverse recursive mining on other functions on the call path, it is possible to gradually mine the vulnerabilities of a large number of related functions and generate vulnerability verification scripts on one or several call paths. This not only ensures the vulnerability mining of the entry function but also further identifies the vulnerabilities existing in other functions, improves the coverage of vulnerability mining, and fully proves that the target software has security issues.
[0047] The present application also provides a vulnerability mining system, an electronic device, and a computer storage medium based on large models and fuzz testing, which possess all the advantages of the above-mentioned vulnerability mining method based on large models and fuzz testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 It is a schematic flowchart of a vulnerability mining method based on large models and fuzz testing according to the present application;
[0050] Figure 2 It is a vulnerability verification script added by the official to verify whether there are still vulnerabilities after patching the vulnerability function.
[0051] Figure 3 It is a reachable path graph from the entry function to the vulnerability function.
[0052] Figure 4 It is a control flow graph of the example function SpriteBuilder().
[0053] Figure 5 It is a schematic diagram of the Prompt template used in the embodiments of the present application.
[0054] Figure 6 It is a schematic diagram of the driving function generated by the large model for the first time.
[0055] Figure 7 It is a schematic diagram of the error reported by the first generated driving function in the compiler.
[0056] Figure 8 It is a schematic diagram of the driving function generated by the large model after the first round of feedback optimization.
[0057] Figure 9 It is a schematic diagram of the driving function generated by the large model after the second round of feedback optimization.
[0058] Figure 10 It is a schematic diagram of the source code of the example call function doNormalize().
[0059] Figure 11 It is a schematic diagram of the constraint conditions analyzed by the large model and the list of parameters to be mutated.
[0060] Figure 12 It is a schematic diagram of the fuzzing driver generated after fixing the parameters.
[0061] Figure 13 It is a schematic diagram of a vulnerability mining system based on a large model and fuzz testing in the present application. Detailed implementation manners
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0063] Accordingly, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0064] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0065] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is usually placed during use. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0066] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0067] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "connected" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0068] Threats from the open-source software supply chain have prompted developers to adopt means such as software composition analysis to detect vulnerabilities in software, and use methods such as manual writing or fuzz testing to construct vulnerability verification scripts to mine for vulnerabilities and thus prove the existence of vulnerabilities. However, these means have the following deficiencies:
[0069] (1) Frequent false alarms: The existing tools have a relatively coarse analysis granularity and cannot accurately determine whether a vulnerable function truly affects downstream projects. This imprecise analysis will generate a large number of false alarms, which are likely to confuse developers' judgment and may prompt them to perform unnecessary dependency upgrades, thereby triggering other dependency conflicts or incompatibility issues.
[0070] (2) Complex writing of vulnerability verification scripts: Due to the high false positive rate of vulnerability detection tools, many developers do not trust the reports provided by these tools. They usually need to rely on specific vulnerability verification scripts to prove the actual impact of vulnerabilities on their own software, so as to evaluate the severity of the vulnerabilities and the necessity of modifying the software. However, manually writing these scripts often requires a deep understanding of the code context, so this process is both complex and time-consuming. In addition, the quality of the scripts highly depends on the experience and ability of the writers, resulting in difficulty in standardizing the writing work, thus increasing the complexity and challenge of vulnerability mining and verification.
[0071] (3) Low efficiency of single fuzz testing: Relying solely on fuzz testing to find vulnerabilities in downstream projects originating from upstream components is not efficient. Fuzz testing lacks effective guidance from input to vulnerability triggering, and arbitrary inputs may not be able to bypass certain specific conditions required to trigger vulnerabilities, thus reducing the efficiency of mining and verifying vulnerabilities.
[0072] Given the significant potential demonstrated by large models in code extraction and generation, the fact that developers of upstream projects often provide vulnerability verification scripts to verify the existence of vulnerabilities in the current version when fixing vulnerable functions, and the advantages of fuzz testing in providing a large number of input samples and vulnerability mining, this application proposes a solution that combines large models and fuzz testing. That is, use the large model to analyze the vulnerability verification scripts and function call information of vulnerable functions, extract the vulnerability verification inputs and the constraint conditions for triggering vulnerabilities, and then guide the mutation direction of fuzz testing inputs. Finally, taking the entry function as the target, mine the vulnerabilities of each function on the vulnerability call path and generate corresponding vulnerability verification scripts, so as to solve the vulnerability reachability and exploitability problems in the target software and achieve the purpose of maintaining and promoting the ecological security of the open-source software supply chain.
[0073] The following further describes this application in detail with reference to the embodiments and the accompanying drawings. As Figure 1 shown, it is the first process schematic diagram of a vulnerability mining method based on large models and fuzz testing in this application, which may include:
[0074] S1, perform static analysis on the target software, and forwardly obtain all call paths from the entry function to the vulnerable function in combination with the entry function information and the vulnerable function information.
[0075] In practical applications, static analysis techniques such as Clang Static Analyzer, Ghidra, or Fortify can be used to analyze the target software. By scanning the target software and combining the entry points with information on known vulnerable functions, a call graph can be obtained. Based on this call graph, starting from the entry function and ending at the vulnerable function, a graph traversal algorithm is used to obtain the set of all function nodes and call edges from the start point to the end point, thereby obtaining all call paths in the forward direction. This method helps to quickly and accurately find all vulnerable call paths and eliminate redundant information unrelated to the actual call relationships.
[0076] S2. Calculate and sort the complexity of all call paths to obtain a queue of call paths; starting from the lowest to the highest complexity of the call paths, successively execute steps S3 to S7 on the call paths in the call path queue.
[0077] In practical applications, a metric standard for path complexity can be defined, such as path length, the number of function calls involved, the number of conditional branches, etc. Calculate the complexity of each path and form a call path queue in descending order.
[0078] As an implementation manner of this application, the following method can be adopted:
[0079] (1) Traverse each call path and generate a control flow graph for each called function. Starting from the entry function, successively count the number of nodes, the number of branch paths, the number of loop structures, the number of nested loops, and the maximum nested loop depth between the current called function and the next called function on the call path.
[0080] (2) According to the number of nodes and the number of branch paths, calculate the cyclomatic complexity between two called functions using the point-edge calculation method, and perform a weighted sum with the complexity generated by the loop structures to obtain the code complexity from the current function to the next called function in the call chain. For the convenience of description, it is hereinafter referred to as function complexity.
[0081] The calculation formula for function complexity is:
[0082]
[0083] Where is the complexity of function , and are the weights of the cyclomatic complexity and the loop structure respectively, is the number of edges in the control flow graph, is the number of nodes in the control flow graph, , , They are the number of loop structures, the number of nested loop structures, and the maximum nesting depth, respectively.
[0084] (3) Sum up the complexities of all called functions and calculate the arithmetic mean. The calculation formula is:
[0085]
[0086] Among them, is the arithmetic mean of the function complexity, is the number of called functions.
[0087] (4) Then perform a weighted sum with the number of functions in each call path to calculate the complexity of each call path, and sort them in ascending order from low to high.
[0088] The calculation formula for the path complexity is:
[0089]
[0090] Among them, is the call path of the complexity, is the weight to balance the function complexity and the path length.
[0091] If there are cases where the complexities of the call paths are the same, then the call path with fewer called functions in the call path is ranked first. If the number of called functions in the call path is the same, then any one of the call paths is ranked first. The calculation and sorting of the complexity of the call path are only performed once.
[0092] The above method combines three main elements: the function cyclomatic complexity metric commonly used in software analysis, the loop-related structure information in the function, and the number of functions in the call path. These three elements consider both the impact of the number of branches and loop structures in the function on the complexity of a single function and the impact of the number of functions on the complexity of the entire path. By calculating and sorting the complexity of each call path in ascending order, in subsequent steps, the call path with the lowest complexity in the current ordered queue can be taken out successively for analysis, which is beneficial to greatly improving the efficiency of the overall process and the accuracy of vulnerability mining.
[0093] (5) If the vulnerability verification script for the entry function has not been generated and there are still un-traversed call paths in the queue, then successively take out the path ranked first in the current queue, that is, the path with the lowest complexity, for subsequent analysis.
[0094] S3, input the Prompt template and the call scenarios of the called functions in the call path into the large model to guide the generation of a driver function that can run the called functions.
[0095] It should be noted that once the call path is determined, all functions that appear on the call path are regarded as calling functions. These calling functions may include library functions and custom functions, etc., which together constitute a complete call chain from the entry function to the vulnerability function. In practical applications, a large model can be used to generate a driving function that can trigger a specific function call scenario, and a Prompt template is designed, which contains the call context and expected behavior of each calling function in the call path. The Prompt template and the call scenario are input into the large model to generate the driving function code.
[0096] In practical applications, the Prompt template can be specifically designed to guide the large model to accurately generate the driving function. In this embodiment, the Prompt template follows the basic principles of being goal-oriented, simple and direct, information-complete, and logically clear in Prompt engineering, and combines the design method of the chain of thought. This design pattern imitates the basic idea of humans when solving the same problem, splitting a large problem into several small steps in sequence, so as to guide the large model to think step by step and improve the generation effect of the model. This design pattern can also clearly output the analysis results of the large model for each small step, thereby reducing the probability of the "hallucination" phenomenon of the large model. In this embodiment, the definition of the calling function is a function that explicitly calls the vulnerability function in the method body in the current call path. The definition of the usage scenario of the calling function is: the parameter format, data structure, and other information passed in when the calling function is used by other functions in the call path. By combining the Prompt template with the usage scenario of the calling function, the large model can fully understand the expected goal of the task and the usage method of the calling function, and then try to generate a driving function that can run the calling function.
[0097] The information input into the large model includes the usage scenario of the calling function in the call path, which can be used as a powerful reference by the large model to help it better construct the driving function for running the calling function.
[0098] S4. If the driving function cannot run the calling function normally, within the limited number of optimization times, the running result is fed back to the large model, and the large model is used to optimize the driving function.
[0099] To ensure that the driving function can run correctly and call the target function, if the driving function cannot run normally, the error information during runtime can be collected, and within the limited number of iterations, these error information are fed back to the large model to guide the large model to correct and optimize the driving function. This automated repair method will help reduce manual intervention and thus improve the running efficiency of the overall process.
[0100] S5. Combine with the driver function, input the vulnerability verification script of the existing vulnerability function and the Prompt template information into the large model, and use the large model to analyze the constraint conditions to select the input parameters that need to be mutated.
[0101] It is possible to identify and select the input parameters that are crucial for vulnerability exploitation for mutation. Combine the driver function and the known vulnerability verification script, input the relevant information into the large model, analyze the constraint relationships between the parameters, and determine which parameters are potential attack surfaces and need to be mutated for testing.
[0102] In practical applications, it is possible to first collect the driver function directly generated by S3 or generated after optimization by S4, as well as the vulnerability verification script of the vulnerability function, and design a Prompt template with a basic idea consistent with the template in S3. The difference is that the main role of this template is to guide the large model to first refer to information such as the calling function, driver function, and vulnerability verification script, then analyze the constraint conditions that need to be satisfied for calling the vulnerability function, and finally select the input parameters that are helpful for triggering the vulnerability and need to be mutated in the subsequent fuzz testing steps, and at the same time select the parameters that are irrelevant to triggering the vulnerability and therefore need to remain unchanged. Furthermore, in subsequent steps, mine the vulnerabilities in the calling function based on the selected parameters.
[0103] S6. Generate seeds, and based on the input parameters that need to be mutated, conduct targeted fuzz testing on the calling function to mine vulnerabilities and obtain the vulnerability verification script of the calling function.
[0104] Generate initial seeds (legal or distorted input data), and for the selected mutated parameters, use the fuzz testing framework to conduct targeted testing on the calling function, attempt to trigger the vulnerability, and record the test results to generate the vulnerability verification script.
[0105] In practical applications, to mine vulnerabilities and obtain the vulnerability verification script of the calling function, the following method can be adopted:
[0106] First, set the parameters that need to be mutated to the specific input parameters selected by the large model in step S5, and fix the parameters that need to remain unchanged. Then generate an initial seed set, and then conduct targeted fuzz testing on the calling function. To encourage the mutation to approach the vulnerability function, the distance between the nodes that the seed can cover and the vulnerability function will be calculated during fuzz testing, and the seeds with a shorter distance will be given a higher priority. Fuzz testing will continue to mine the vulnerability until the end condition is met: exceeding the limited running time or having mined the input that can verify the existence of a vulnerability in the calling function. Furthermore, the current input can be combined with the driver function to obtain the corresponding vulnerability verification script.
[0107] In this embodiment, a directed grey-box fuzzer is used to perform fuzz testing on the calling function. By calculating the distance between the seed-coverable nodes and the vulnerable function, and taking the closer to the vulnerable function as the better standard, the fuzz testing can move towards the vulnerable function, thereby attempting to discover the vulnerabilities of the calling function. Compared with coverage-based fuzz testing, directed fuzz testing preferentially retains the seeds whose coverable nodes are closer to the vulnerable function, so the efficiency of vulnerability discovery can be guaranteed within a limited time.
[0108] S7. Use the vulnerability verification script of the calling function as a new reference, and recursively reverse-mine the vulnerabilities of the calling functions on the current call path until the vulnerabilities of the entry function are discovered.
[0109] Vulnerabilities can be recursively mined upstream along the call chain until the entry function is reached. Use the newly discovered vulnerability verification script as new reference information, and repeat steps S3 to S6, but this time for each calling function on the current call path, reverse-mine its potential vulnerabilities until the vulnerabilities of the entry function are finally discovered.
[0110] In practical applications, to recursively reverse-mine the vulnerabilities of the calling functions on the current call path until the vulnerabilities of the entry function are discovered, the following method can be specifically adopted:
[0111] (1) When the vulnerability verification script obtained in step S6 can prove that the calling function has vulnerabilities, use this script as a new reference and regard the current calling function as a new vulnerable function.
[0112] (2) Reverse-select the next function immediately following in the call path as the calling function, and repeat steps S3 to S6 for it to generate the corresponding vulnerability verification script;
[0113] (3) If the optimization task is not completed in step S4 or the directed fuzz testing in step S6 exceeds the limited running time, it means that the vulnerabilities of the current calling function cannot be discovered. The current call path needs to be marked as infeasible and restarted from step S2.
[0114] (4) Repeat (1) to (3) in step S7, with the entry function as the end point, continuously and recursively mine the vulnerabilities of each calling function and generate vulnerability verification scripts.
[0115] (5) If the vulnerabilities of the entry function are finally successfully discovered, it means that the target software has at least one call path that can trigger vulnerabilities at the entry point, thus verifying the reachability and exploitability of the vulnerabilities, and further proving that the target software has a relatively high security risk.
[0116] During the process of reverse generating vulnerability verification scripts for other functions, when a vulnerability is discovered in a calling function, the generated vulnerability verification script is used as a new reference, the current calling function is regarded as the vulnerable function, and the next immediately following function is selected reversely from the call path as the calling function and the vulnerability mining process is restarted. If the vulnerability of a certain function cannot be discovered, a new path is selected from the call path queue for traversal, and this process continues until the vulnerability of the entry function is successfully discovered or the call path set is empty.
[0117] This application first obtains all call paths from the entry function to the vulnerable function of the target software based on static analysis technology, sorts them according to the complexity of the paths, and forms an ordered call path queue. Then, the path with the lowest complexity in the queue is taken out successively for traversal, and a driver function for running the vulnerable function is generated using a large model for effective fuzz testing. Then, the vulnerability verification script of the vulnerable function and the driver function are input into the large model, and it is used to analyze the constraint conditions of the calling function, and then the parameters that need to be mutated and those that do not need to be mutated are selected. After that, this application uses the directed fuzz testing technology to mine the vulnerability of the calling function based on the selected parameters, and uses the distance between the seed-covered nodes and the vulnerable function as a criterion to guide the parameter mutation process, that is, the higher the priority of the seed with the closer distance. When the vulnerability of the calling function is successfully discovered, the fuzz testing input and the driver function are combined to generate a new vulnerability verification reference script, and the vulnerability of other calling functions is continuously mined recursively in reverse and the vulnerability verification script is generated until the entry function is reached. This application aims to improve the accuracy and efficiency of vulnerability mining and verify whether the target project is affected by specific vulnerabilities, so as to solve the problems of vulnerability reachability and exploitability, and provide a strong guarantee for the security of the open source software supply chain.
[0118] Figure 2 A vulnerability verification script provided for the developers of the Apache Commons IO project for the CVE-2021-29425 vulnerability reported in the project. This script can run smoothly in the patched version, but will report an error in the vulnerable version, and it can be used to mine whether there are vulnerabilities in other functions that call the vulnerable function.
[0119] Taking the Apache Commons IO 2.6 version and the SmartSprites project that uses this vulnerable version as an example, the vulnerability mining method based on large models and fuzz testing technology proposed in this application is described.
[0120] Step 1: First, use the widely used Soot static analysis framework in Java to perform static analysis on the target project SmartSprites to obtain the function call graph between this project and the vulnerable project Apache Commons IO. Then, starting from the entry function of the target project and ending with the vulnerable function getPrefixLength() in the vulnerable project, perform a depth-first traversal operation, and successfully obtain Figure 3 two call paths in
[0121] . Each node in the graph is represented in the form of a Soot signature, and the basic structure is <class name: function return value function name (function input parameters)>. Figure 4 shows the control flow graph of the third node SpriteBuilder() starting from the entry function in the call path, and marks the path from the entry of SpriteBuilder() to reWriteCssFiles(). It can be seen from the figure that the directed maximum connected subgraph between the two functions has 32 edges, 32 nodes, 1 loop structure, 0 levels of nesting and the maximum nesting depth is 0. We take the weight in the function complexity calculation formula as 0.7, take as 0.3, and the complexity of SpriteBuilder() can be calculated as 1.7 according to the following formula.
[0122]
[0123] Set the left path in Figure 3 as path 1 and the right path as path 2, calculate the complexity of other functions on the two call paths, sum them up and take the arithmetic mean, and the average function complexity of path 1 is about 1.35, while the average function complexity of path 2 is about 1.7. Then, according to the path complexity calculation formula , take the weight W as 0.1 and calculate the complexity of the two paths respectively. There are 11 functions on path 1 and 12 functions on path 2. Thus, the complexity of path 1 is calculated as 2.45, and the complexity of path 2 is 2.9. Therefore, the priority of path 1 is higher than that of path 2, thus forming an ordered call path sequence. Then, first take out path 1 for the subsequent steps.
[0124] Step 3: Combine a specially designed Prompt template with the call scenarios used by other functions on the function path and provide it to the large model, and use the abstraction and summarization ability of the large model to try to generate a driving function that can run the call function. Figure 5It is a Prompt template presented in Markdown format. Since large language models have much stronger understanding and analysis capabilities for English than for Chinese, we write the Prompt template in English. Here, we introduce each element in the thought chain that makes up the template. The objective function in the template is the calling function that needs to generate a driver function, and the reference function is the corresponding calling scenario:
[0125] (1)Overall analysis
[0126] Based on the content provided at the beginning of the template, identify the task that needs to be completed; based on the <targetmethod>< / targetmethod> (2)Analysis of the relationship between the calling function and the called function
[0127] (3)Reference function analysis
[0128] Indicate which packages should be imported in the driver function to avoid the "symbol cannot be resolved" error; indicate what inputs are used in the reference function to call the objective function; fully consider how these inputs should be constructed, especially when these inputs are instances of custom classes; (4)Access modifier analysis
[0129] (5)Ask the large language model to write a driver function that includes all necessary elements and can be directly run in the original project.
[0130] Figure 6 Taking doNormalize() as the objective function and its normalize() as the reference calling scenario, the driver function obtained by inputting it into the large language model in combination with the Prompt template. It is not difficult to see from the generated result that the large language model fully understood the actual function of the objective function, successfully constructed all necessary inputs, and then generated a driver function with the correct structure that calls the objective function. Furthermore, the large language model also successfully analyzed that the objective function is a private function and used reflection technology to call it.
[0131] Step 4: Test the driver function generated in Step 3. As Figure 7 shown, we found a symbol not found error, so we fed the error message back to the large language model and asked it to optimize the driver function. The optimization result is as Figure 8 shown. The large language model successfully added the missing packages to the driver function. Although the driver function can already run, the compiler prompts a missing package declaration, so we feed the prompt message back to the large language model again and finally obtain the Figure 9 shown driver function. After verification, the current driver function has no errors and can run normally and call the objective function.
[0132] Step 5: Combine the optimized driver function, the vulnerability verification script of the vulnerability function, and a Prompt template with the same construction idea as the template in Step 3, input them into the large model, and require it to analyze the constraints required to call the vulnerability function, and then select the input parameters that need to be mutated during the fuzz testing process, as well as the parameters that are irrelevant to vulnerability mining and thus need to remain unchanged. Figure 10 The source code of the calling function is doNormalize(), and the vulnerability function is getPrefixLength(). The constraint conditions extracted by the large model, as well as the selected parameters that need to be mutated and the parameters that need to remain unchanged, are as Figure 11 shown.
[0133] Step 6: Fix these parameters in the driver function according to the parameters that remain unchanged selected by the large model in Step 5 to obtain the fuzz driver, as Figure 12 shown. For the parameters that need to be mutated, such as filename, if it has a certain mutation range, randomly generate an initial value within the range. Otherwise, we initialize and generate an initial seed for it using a dictionary and a template based on its type, as the initial seed set.
[0134] After assigning energy to each seed in the seed set, mutate and execute it to generate the corresponding test results and obtain a new seed set. The seed execution paths obtained from the test results can be used to calculate the priority for the new seed set. We use an adaptive strategy to sort the priorities of the seeds. Let v be the corresponding vulnerability path, be the path that does not cover v but is the closest to v. The forking path condition is c, and the forking conditions can be the following: true, false, , , , , , , where a and b are variables or constants. We define a distance function to sort the priorities of the seeds, and the distance function is as follows:
[0135]
[0136] Among them, is a constant representing the minimum distance. It can be considered that the smaller the value, the closer the branch path is to the vulnerability path. Furthermore, the distance metric calculation can be used to sort the seed set by priority. The seeds with branch paths closer to the vulnerability path during execution are given higher priorities. The higher the priority of a seed, the higher the energy assigned to it during mutation, and the greater the probability of mutating a seed that can uncover a vulnerability.
[0137] The fuzz testing process continues to execute until the program meets the end condition: reaching the limited running duration or generating inputs that can prove vulnerabilities in the called functions. Then, the testing ends, and the inputs are combined with the driver script to generate corresponding vulnerability verification scripts.
[0138] Step 7: If a verification script that can uncover vulnerabilities in the called function is successfully generated and the current call path has not been fully traversed, use this script as a new reference, consider the current called function as a vulnerable function, and take the next function immediately following on the call path as the called function. Continue to recursively reverse-mine for vulnerabilities in other functions and generate vulnerability verification scripts in a loop until vulnerabilities in the entry function are mined, thereby proving the reachability and exploitability of the vulnerabilities, and further confirming that there are security issues in the current project.
[0139] Through the above method, the present application realizes a vulnerability mining method based on large models and fuzz testing technology, providing a faster and more accurate solution for the reachability and exploitability problems of vulnerabilities in the open-source software supply chain ecosystem.
[0140] As Figure 13 shown, it is a schematic diagram of a vulnerability mining system based on large models and fuzz testing according to the present application, which may include:
[0141] A static analysis module for statically analyzing the target software, combining the entry function information and the vulnerable function information, and obtaining all call paths from the entry function to the vulnerable function in a forward direction;
[0142] A mining module for calculating and sorting the complexity of all call paths to obtain a set of call path queues; successively execute the following steps for the call paths in the call path queue from the lowest to the highest complexity of the call paths:
[0143] Input the Prompt template and the call scenario of the called function in the call path into the large model to guide the generation of a driver function that can run the called function;
[0144] If the driver function cannot run the called function normally, within the limited number of optimization times, feedback the running result to the large model and optimize the driver function through the large model;
[0145] Combine the driver function, input the vulnerability verification script of the existing vulnerable function and the Prompt template information into the large model, and use the large model to analyze the constraint conditions and select the input parameters that need to be mutated;
[0146] Generate seeds and perform targeted fuzz testing on the called function based on the input parameters that need to be mutated to mine vulnerabilities and obtain the vulnerability verification script of the called function;
[0147] Take the vulnerability verification script of the calling function as a new reference, and recursively reverse-mine the vulnerabilities of the calling functions on the current call path until the vulnerabilities of the entry function are mined.
[0148] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules can be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] In addition, in each embodiment of the present invention, the modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0150] The embodiments of the present application also provide an electronic device, which may include one or more processors, a memory, and a communication interface.
[0151] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.
[0152] Among them, the communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the steps of the vulnerability mining method based on the large model and fuzz testing in the embodiments of the present application.
[0153] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or any combination of other programmable logic devices, transistor logic devices, and hardware components. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. The processor can also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device to execute the method steps provided in the above embodiments.
[0154] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, Figure 6 only a thick line is used to represent it in [description], but it does not mean that there is only one bus or one type of bus.
[0155] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, it implements the steps of the vulnerability mining method based on a large model and fuzz testing described in any of the above embodiments.
[0156] The computer-readable storage medium involved in the present application includes a Random Access Memory (RAM), a memory, a Read-Only Memory (ROM), an Electrically Programmable ROM, an Electrically Erasable Programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium well-known in the technical field.
[0157] For the description of the relevant parts in the vulnerability mining system, electronic device, and computer-readable storage medium based on a large model and fuzz testing provided by the embodiments of the present application, please refer to the corresponding detailed description in the vulnerability mining method based on a large model and fuzz testing provided by the embodiments of the present application, and will not be elaborated here. In addition, the parts of the above technical solutions provided by the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0158] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A vulnerability mining method based on large models and fuzz testing, characterized in that, Including: S1. Perform static analysis on the target software, combine the entry function information and vulnerability function information, and obtain all call paths from the entry function to the vulnerability function in a forward direction; S2. Calculate and sort the complexity of all call paths to obtain a set of call path queues; starting from the lowest to the highest complexity of the call paths, successively execute steps S3 to S7 on the call paths in the call path queue; S3. Input the Prompt template and the call scenario of the calling function in the call path into the large model to guide the generation of a driving function that can run the calling function; S4. If the driving function cannot run the calling function normally, within the limited number of optimization times, feedback the running result to the large model to optimize the driving function through the large model; S5. Combine the driving function, input the vulnerability verification script of the existing vulnerability function and the Prompt template information into the large model, and use the large model to analyze the constraint conditions to select the input parameters that need to be mutated; S6. Generate seeds, and based on the input parameters that need to be mutated, perform directed fuzz testing on the calling function to discover vulnerabilities and obtain the vulnerability verification script of the calling function; The step of discovering vulnerabilities and obtaining the vulnerability verification script of the calling function includes: Set the parameter that needs to be mutated as the input parameter that needs to be mutated, and fix the parameters that remain unchanged; generate an initial seed set, and perform directed fuzz testing on the calling function according to the initial seed set to obtain the vulnerability verification script of the calling function: In fuzz testing, calculate the distance between the nodes that the seeds can cover and the vulnerability function, and assign higher priority to the seeds with shorter distances; continuously perform fuzz testing to discover vulnerabilities until the end condition is met: exceeding the limited running duration or discovering an input that can verify the existence of vulnerabilities in the calling function; combine the current input with the driving function to obtain the corresponding vulnerability verification script; S7. Use the vulnerability verification script of the calling function as a new reference, and regard the current calling function as a new vulnerability function; reverse-select the next adjacent function from the call path as the calling function, and repeat steps S3 to S6 to generate the corresponding vulnerability verification script; taking the entry function as the end point, continuously loop and recursively discover the vulnerability verification scripts of each calling function until the vulnerability of the entry function is discovered.
2. The vulnerability mining method based on large models and fuzz testing according to claim 1, wherein, The step of obtaining all call paths from the entry function to the vulnerability function in a forward direction includes: Use static analysis technology to analyze the target software to obtain a call graph. Based on the call graph, taking the entry function as the starting point and the vulnerability function as the end point, use the graph traversal algorithm to obtain the set of all function nodes and call edges from the starting point to the end point, and obtain all call paths in a forward direction.
3. The vulnerability mining method based on large models and fuzz testing according to claim 2, characterized in that, The step of calculating and sorting the complexity of all call paths to obtain a set of call path queues includes: Traverse each call path and generate the control flow graph of each calling function. Starting from the entry function, successively count the number of nodes, the number of branch paths, the number of loop structures, the number of nested loops, and the maximum nested loop depth from the current calling function to the next adjacent calling function on the call path, and record them as inter-function parameters; According to the control flow graph and inter - function parameters, calculate the cyclomatic complexity between two calling functions using the point - edge calculation method, and perform a weighted sum with the complexity generated by the loop structure to obtain the code complexity from the current function to the next adjacent calling function. Sum the complexities of all calling functions and calculate the arithmetic mean. Perform a weighted sum of the arithmetic mean and the number of functions in each calling path to calculate the complexity of each calling path.
4. The vulnerability mining method based on large models and fuzz testing according to claim 3, characterized in that, The code complexity from the current function to the next adjacent calling function is obtained through the following formula: Among them, is the complexity of the function , and are the cyclomatic complexity and the weight of the loop structure respectively, is the number of edges in the control flow graph, is the number of nodes in the control flow graph, , , are the number of loop structures, the number of nested loop structures, and the maximum nesting depth respectively.
5. A vulnerability mining system based on large models and fuzz testing, characterized in that, Including: A static analysis module for statically analyzing the target software. Combining the entry function information and vulnerability function information, it obtains all calling paths from the entry function to the vulnerability function in a forward manner. A mining module for calculating and sorting the complexities of all calling paths to obtain a queue of calling paths. Starting from the lowest - complexity calling path in the queue of calling paths, successively perform the following steps (1) to (5) on the calling paths: (1) Input the Prompt template and the calling scenario of the calling function in the calling path into the large model to guide the generation of a driver function that can run the calling function. (2) If the driver function cannot run the calling function normally, within the limited number of optimization times, feedback the running result to the large model to optimize the driver function through the large model. (3) Combine the driver function and input the vulnerability verification script of the existing vulnerability function and the Prompt template information into the large model. Use the large model to analyze the constraint conditions and select the input parameters that need to be mutated. (4) Generate seeds and perform directed fuzz testing on the calling function based on the input parameters that need to be mutated to dig out vulnerabilities and obtain the vulnerability verification script of the calling function. The process of digging out vulnerabilities and obtaining the vulnerability verification script of the calling function includes: Set the parameters that need to be mutated as the input parameters that need to be mutated and fix the parameters that remain unchanged; generate an initial seed set and perform directed fuzz testing on the calling function according to the initial seed set to obtain the vulnerability verification script of the calling function: During fuzz testing, calculate the distance between the nodes that the seeds can cover and the vulnerability function, and assign higher priority to the seeds with shorter distances; continuously perform fuzz testing to dig out vulnerabilities until the end condition is met: exceeding the limited running time or digging out an input that can verify the existence of a vulnerability in the calling function; combine the current input with the driver function to obtain the corresponding vulnerability verification script. (5) Use the vulnerability verification script of the calling function as a new reference and regard the current calling function as a new vulnerability function; reverse - select the next adjacent function in the calling path as the calling function and repeat steps (1) to (4) to generate the corresponding vulnerability verification script; taking the entry function as the end point, continuously loop and recursively dig out the vulnerability verification scripts of each calling function until the vulnerability of the entry function is dug out.
6. An electronic device, characterized in that, Including: A memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the vulnerability mining method based on large models and fuzz testing as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the vulnerability mining method based on large models and fuzz testing as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Method and system for carrying out vulnerability repair verification based on large model
CN118277144A
Fuzzy testing a software system
EP4006733A1