Method for setting up fuzzing test for fuzz target
By inputting fuzz target documentation into a language understanding AI to generate API calls and arguments, the method addresses the challenge of creating fuzz drivers without a learning base, facilitating efficient fuzz testing.
Patent Information
- Application Number
- JP2024220487
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-17
- Publication Date
- 2025-07-01
AI Technical Summary
Existing technologies face challenges in generating high-quality fuzz drivers without a learning base, such as existing test cases, limiting effective fuzz testing.
A method involving inputting fuzz target documentation into a language understanding AI to generate API calls and arguments, and using these to create fuzz drivers, with optional feedback loops for improvement.
Enables effective fuzz testing by automatically generating fuzz harnesses even when only documentation is available, determining API call sequences and handling arguments effectively.
Smart Images

Figure 2025097950000001_ABST
Abstract
Description
Technical Field
[0001] Fuzzing is more widely adopted as a dynamic software testing method. Fuzzing or fuzz testing is an automated software testing technique that generally involves providing invalid, unexpected, or random data as input to a fuzz target, typically in the form of a computer program.
Background Art
[0002] A fuzz driver connects a fuzzer to a fuzz target and injects the input generated by the fuzzer into the software under test. To generate a fuzz driver, there must be some learning base. Such a base can be existing test cases that can record program traces or existing unit tests, which can be converted into fuzz drivers. Without such a learning base, it may not be possible to generate a fuzz driver, or if generated, the quality will be low.
Summary of the Invention
[0003] A first aspect of the present disclosure relates to a method for generating a fuzz driver for fuzz setup. The method of the present disclosure is - inputting the documentation of the fuzz target into a language understanding artificial intelligence (AI); - generating application programming interface (API) calls and their arguments from the documentation by the language understanding AI; - generating at least one fuzz driver from the API calls and their arguments and includes.
[0004] A second aspect of the present disclosure relates to a method for training a language understanding AI for generating a fuzz driver for fuzz setup. The method of the present disclosure is - The step of generating a fuzz driver for fuzz setup according to claim 1 of the first aspect of the present disclosure; - The step of selecting and executing the fuzz driver according to claim 7 of the first aspect of the present disclosure; - The step of generating a fitness value of the fuzz driver when the fuzz driver is executed; - The step of feeding back the fitness value to the language understanding AI; - The step of updating the language understanding AI using the fitness value and including.
[0005] The third aspect of the present disclosure relates to a computer system configured to implement the method according to the first aspect (or an embodiment thereof) and / or the second aspect (or an embodiment thereof) of the present disclosure.
[0006] The fourth aspect of the present disclosure relates to a computer program configured to implement the method according to the first aspect (or an embodiment thereof) and / or the second aspect (or an embodiment thereof) of the present disclosure.
[0007] The fifth aspect of the present disclosure relates to a computer-readable medium or signal storing and / or including the computer program according to the fourth aspect (or an embodiment thereof) of the present disclosure. The techniques of the first to fifth aspects can have advantageous technical effects.
[0008] The techniques of the present disclosure enable effective generation of fuzz drivers, or in other words, enable fuzz testing by automatic fuzz harness generation. This generation aims at the case where there are no existing test cases. For example, a fuzz driver can be generated when only some documentation exists. Scope statements and product requirement documents are good sources of information as they include all developed interfaces. Additional pre-test cases may improve fuzz driver generation.
[0009] With the techniques of the present disclosure, in the context of automatic fuzzer generation, it is possible to determine which API calls to select, in what order to call the API calls harness, and how to handle the arguments for the pre-selected API calls.
[0010] With the techniques of the present disclosure, it is possible to generate a fuzzer even when there is no source code or when a black box fuzzer or a debugger-driven fuzzer is used. The present invention relates to fuzzing as a dynamic software testing method for enabling fuzz testing, particularly for automatic fuzz harness generation.
[0011] According to the techniques of the present invention, natural language understanding is used to analyze, for example, software documentation and extract in what order and / or with which inputs the application programming interface (API) calls of the software under test should be used. Using these API call sequences and parameters, a high-quality fuzzer can be generated, that is, it is possible to cover most of the program code during fuzzing.
[0012] In this specification, several terms are used as follows. The term "fuzzing or fuzz test" refers to an automated process that sends generated invalid, unexpected, or random inputs to a target such as a computer program and monitors the target's reaction to exceptions such as crashes, failures of built-in code assertions, and potential memory leaks.
[0013] A "fuzzer or fuzzing engine" generates semi-valid inputs that are valid enough not to be directly rejected by the parser for the fuzzing target, but are "invalid enough" to produce unexpected behavior deeper in the program and expose corner cases that are not properly handled.
[0014] A "fuzz target" is a software program or function intended to be tested by fuzzing. The main feature of a fuzz target is that it can consume inputs such as binary, library, API, or bytes for example.
[0015] A "glue code, wrapper, harness, or fuzz driver" connects the fuzzer to the fuzz target and injects the inputs generated by the fuzzer into the software under test.
[0016] A "fuzz test" is a composite version of the fuzzer and the fuzz target. At this time, the fuzz target may be instrumented code to which the fuzzer is attached for its input. The fuzz test is executable. The fuzzer can also start, monitor, and stop multiple execution fuzz tests (usually hundreds or thousands per second) using slightly different inputs generated by the fuzzer respectively.
[0017] A "test case or input" is one specific input and test execution from a fuzz test. Usually, for reproducibility, interesting executions (detection of new code paths or crashes) are saved. In this way, a specific test case, and its corresponding input, can also be executed on a fuzz target that is not connected to the fuzzer, i.e., its release version.
[0018] "Instrumentation" is used to make coverage metrics monitorable, for example during compilation. Instrumentation is the injection of instructions into a program to generate feedback from execution. This is mainly achieved by the compiler and can describe, for example, the code blocks reached during execution.
[0019] "Coverage-guided fuzzing" refers to using code coverage information as feedback during fuzzing to detect whether the input has caused the execution of a new code path / block.
[0020] "Mutation-based fuzzing" refers to creating new inputs by randomly mutating a set of known inputs (corpus). "Generation-based fuzzing" refers to creating new inputs from scratch, for example, by using an input model or input grammar.
[0021] A "mutator" is a function that takes bytes as input and outputs small random mutations of the input. A "corpus" (plural: corpora) is a set of inputs. The initial inputs are seeds.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0023] Figure 1 discloses and proposes a method for generating a fuzzer driver for fuzz set-up. The method proposed in this disclosure uses natural language understanding to analyze some fuzz target documentation, generate application programming interface (API) calls and their arguments, and generate at least one fuzz driver for targeting the fuzz target.
[0024] The fuzz target may include software that is designed to control, adjust, and / or monitor a cyber-physical system, especially as at least one computing unit of a vehicle, such as an automobile, a robot, an IoT device, an electric assist bicycle, a Car-2-X infrastructure, and especially a vehicle. In particular, the software may be embedded software designed to run on an embedded (i.e., task-specific) system. The method of this disclosure may be related to automated security testing.
[0025] The method proposed in this disclosure may be used during the development before the release of the software under test. Proprietary software or known software may be fuzzed when flashed to the target so as to have a system that can be actually tested rather than a simulated system. Also, unknown software distributions or fuzz software distributions from third parties that may have a black box format (i.e., as binary or shared library files) may be fuzzed.
[0026] Furthermore, the method proposed in this disclosure can determine in what order and with what inputs the application programming interface (API) calls of the software under test should be used.
[0027] The first step of the method proposed in this disclosure is to input the documentation of the fuzz target into a language understanding artificial intelligence (AI) (11). The documentation may include at least one of a handbook, a scope statement, a product requirements document, code, binary debug symbols, logs such as communication records between the fuzz target, and / or, for example, programming comments from source code level work. The language understanding AI may comprise at least one of a large language model (LLM), natural language processing (NLP), natural language understanding (NLU), and / or the language understanding AI is a trained model or a base model.
[0028] Optionally, the source code of the fuzz target can be input into the language understanding AI. Such input improves the quality of the generated API calls and their arguments, and, along with it, also improves the quality of the generated fuzz driver. Without inputting the source code, it is also possible to fuzz the fuzz target using a black box fuzzer or a debugger-driven fuzzer.
[0029] The second step is to generate application programming interface (API) calls and their arguments from the documentation by the language understanding AI (12). The arguments are parameters for each API, such as, for example, the path of function parameters. The arguments vary according to each API.
[0030] The output of the generation may be saved in a list or a set, and such a list or set may be saved together with optional fitness values such as the following. {{API1(arg1,arg2,…,arg m ),value fitness},{API2(arg2,arg3,…,arg k ),value fitness},…{…}} The API call corresponds to API x and the argument is argx corresponds to.
[0031] If there is no optional fitness value (value fitness ), the above list can be used by omitting those fitness values. Arguments may occur multiple times, for example, with different APIs. Further, an API may occur multiple times with various arguments.
[0032] Optionally, when generating API calls and their arguments, a fitness value for each API call may be generated. In a setup without feedback to the AI, the AI estimates how good the fuzz driver to be generated later will be when creating the API call. In a setup with feedback to the AI, the feedback loop allows the AI to create better estimates.
[0033] A fitness function can be defined as a function that takes a candidate solution to a problem as input and generates, as output, how well or how "good" that solution is for the problem under consideration. Here, the fitness value of a fuzz driver may be measured by the generated code coverage of the executed fuzz driver.
[0034] The third step is to generate at least one fuzz driver from the API calls and their arguments (13). The generation of one or more fuzz drivers may be realized by a fuzz driver generator. Optional fitness values may not be used for the generation of fuzz drivers.
[0035] Accordingly, using the generated APIs and arguments, a fuzz driver is generated. Each API and argument may generate no fuzz drivers at all or a large number of fuzz drivers. That is, Gen fuzzdriver (API n ,arg1,…,argm ) -> fuzzdriver f (API n , arg1, …, arg m ), f ∈ N0 and f is within the set of all natural numbers.
[0036] The resulting fuzzdrivers can also be saved along with their corresponding (optional) fitness values. That is, {fuzzdriver f (API n , arg1, …, arg m ), value fitness}, …, {…}} is obtained.
[0037] Optionally, at least one of the generated fuzzdrivers is input into the set of fuzzdrivers and / or the corpus. The corpus can also be used for further systems or methods and also for the training of AI. This approach is particularly applicable, for example, to very broad fuzzdrivers for standard applications.
[0038] Optionally, one or two of the following steps are performed as part of the method. A step of selecting at least one fuzzdriver according to a selection strategy, where the selection strategy is a heuristic or metric. For example, the actual metric for "next best" selection may be calculated from the algorithm or the AI itself. Further, the selection strategy can include at least one of first-in-first-out, last-in-first-out, highest fitness value, deepest abstract syntax tree (AST) derived by static analysis. The fuzzdriver selection strategy may match or conform to the fuzzing process. The selection step may include not only which fuzzdriver is selected but also in which order the fuzzdrivers are selected.
[0039] Next, the execution of the selected fuzz driver for the fuzz target is performed until the termination criteria are met. The termination criteria may include at least one of a new coverage gain over several minutes, a crash being detected, and covering a code location not reached by the previous fuzz driver. This step includes the actual fuzzing of using the selected fuzz driver against the fuzz target. Fuzzing may target only programs, such as compiled code (e.g., regarding c++, C, python, java, javascript, rust, php, etc.) or more general executable code.
[0040] Furthermore, optionally, the fitness value of the fuzz driver is measured or generated based on the generated code coverage of the executed fuzz driver. Thus, the fitness value can indicate to what extent or which areas of the fuzz target's code each fuzz driver reached. The fitness value is then fed back to the language understanding AI for training purposes to improve further generations of API calls and arguments.
[0041] Also optionally, the measured fitness value is fed back into the generation of at least one fuzz driver. Generation may then be re-triggered depending on the fitness of the fuzz driver. Fuzz drivers that cannot be compiled may be marked with a very low fitness value (e.g., 0), but can be retained for further rapid improvement.
[0042] Figure 2 schematically shows a computer system 20 that can generate a fuzz driver for fuzz setup by a language understanding AI using the techniques of the present invention. The computer system 20 is designed to implement the method 10 according to FIG. 1 and the method 40 according to FIG. 4. The computer system 20 can be realized by hardware and / or software. Therefore, the system shown in FIG. 2 can be regarded as a computer program designed to execute or implement the method 10 according to FIG. 1 and the method 40 according to FIG. 4.
[0043] Documentation is provided as an input 21 to the system. This documentation includes the software under test, i.e., the documentation of the fuzz target, as described above with respect to FIG. 1. For example, it may be a scope statement and product requirements document, and comments from source code level work.
[0044] The language understanding AI 22 is input with the documentation. The AI is a trained model for identifying all (or at least the most interesting) API calls and arguments from the documentation. The AI used here may be some large language model (LLM), or other natural language processing (NLP) means, or more specifically natural language understanding (NLU) means.
[0045] A base model can also be used. In that case, the follow-up training can include a specific source code base, such as typical software from the company's cloud and its documentation from, for example, the company's document server.
[0046] The output 23 of the AI is the API calls and arguments, which may be provided as a list or set of the generated (or extracted) API calls and their respective arguments.
[0047] An identified API and arguments for generating a fuzz driver are input to the fuzz driver generator 24. Since the interesting APIs and arguments have already been identified by the AI 22, the fuzz driver generator 24 can remain simple. The fuzz driver generator 24 generates a fuzz driver 25, which is an actual driver (or harness) for injecting the input generated by the fuzzer into the software under test. One or more fuzz drivers 25 are generated by the fuzz driver generator 24.
[0048] Elements of the system 20 mentioned above or parts thereof may correspond to steps of the method 10. In particular, the details of the method steps may apply to the system elements. FIG. 3 schematically shows a computer system 20 that can use the techniques of the present disclosure for generating a fuzz driver. The computer system 20 may correspond to the computer system 20 of FIG. 2. The computer system 20 is adapted to execute the method 10 according to FIG. 1 and the method 40 according to FIG. 4. In particular, the computer system 20 of FIG. 3 is adapted to execute the training method 40 according to FIG. 4. The computer system 20 can be realized in hardware and / or software. Thus, the system shown in FIG. 3 can be regarded as a computer program designed to execute the method 10 according to FIG. 1 and the method 40 according to FIG. 4.
[0049] The elements 21, 22, 23, 24, 25 of the computer system 20 shown in FIG. 3 may correspond to, or be identical to, the elements 21, 22, 23, 24, 25 of the computer system 20 shown in FIG. 2. In other words, at least one of the elements 26, 27, 28, and 29, and their connections, may be optional for the system 20. The system or method according to FIG. 3 may be referred to as performing a fuzz setup or generating a fuzz driver, and applying the fuzz driver or their selection to a fuzz target.
[0050] The fuzz setup 26 takes the fuzz driver 25 as input and executes the fuzz driver 25, that is, it is used to fuzz the fuzz target 27 using the fuzz driver 25. Here, optionally, from the perspective of the coverage that each fuzz driver 25 can generate, the fitness of the fuzz driver 25 can be measured. Thereby, a fitness value is assigned to the fuzz driver 25, or an existing fitness value is updated or modified. The existing fitness value may be assigned to the API call corresponding to the fuzz driver 25 by the AI 22. Since multiple fuzz drivers 25 may be applied to the fuzz target 27, the fuzz driver 25 may be tried or applied for only a short time compared to the overall fuzz campaign.
[0051] The system 20 includes feedback 28 or a feedback loop from the fuzz setup 26 to the AI 22 and / or the fuzz driver generator 24. The feedback returns details from the fuzzing to the AI 22 and / or the fuzz driver generator 24 to improve the API calls and / or the generation of the fuzz drivers 25. The details may include the acceptance of the fuzz driver 25, the fuzzing time, etc. Alternatively or additionally, the feedback information may include the fitness value for each fuzz driver 25.
[0052] Such a fitness value of the fuzz driver 25 is measured or generated by the generated code coverage of the executed fuzz driver 25. Therefore, the fitness value can indicate to what extent or which area of the code of the fuzz target 27 each fuzz driver 25 has reached. The fitness value is then fed back to the language understanding AI 22 for training purposes to improve further generations of API calls and arguments. This can be implemented, for example, as the refinement of the prompt when working with an LLM.
[0053] As an alternative or in addition, the fitness value is fed back to the fuzz driver generator 24. Then, generation may be re-triggered depending on the fitness of the fuzz driver 25. A fuzz driver 25 that cannot be compiled may be marked with a very low fitness value (e.g., 0), but can be retained for further quick improvement.
[0054] Furthermore, the source code 29 is input into the AI 22. This may be a source code repository or any other code storage means. For better results with the generated API calls and arguments, the source code may be the software to be tested, i.e., the code of the fuzz target 27.
[0055] The source code 29 is further input into the fuzz driver generator 24. When the source code becomes available to the fuzz driver generator 24, for example, the function signature can be directly obtained (e.g., copy & paste) and used as a fuzz driver. Unit tests that can be easily converted into fuzz drivers may become available. The source code can reveal the flow of the input through some static analysis (e.g., AST). In such a case, the fuzz driver can inject the input and potentially trigger most of the program.
[0056] The source code 29 is further input into the fuzz setup 26. The source code can reveal whether units, components, and / or systems should execute in a particular context such as some initialization. The source code also provides proper initialization of variables or memory areas. The source code also helps in determining how to (re)set the fuzz target to a known state or whether the detection result is a false positive.
[0057] Figure 4 is a flowchart showing a method 40 for training a language model or a language understanding AI. The language model or AI is set up to automatically generate fuzz drivers for fuzz setups or fuzz tests. The training method 40 uses part of the generation method 10. Although not shown in detail in Figure 4, the training method 40 can also use parts of the generation method 10 that are not described in relation to Figure 4. The training method 40 may also include all parts of the generation method 10. In particular, the training method 40 can include at least one including all dependent claims of the generation method 10.
[0058] As a result, the success of the fuzzing action of the fuzz driver is evaluated and fed back to the AI. In the AI, the feedback is used to adapt the AI, which is done, for example, by adapting the weights to improve the generation of the next API calls generated by the fuzz driver and their arguments.
[0059] A method 40 for training a language understanding AI to generate a fuzz driver for fuzz setup includes the following steps. In a first step, a fuzz driver for fuzz setup is generated (41) according to claim 1 of the training method 10, which includes the following (sub)steps.
[0060] - Inputting the documentation of the fuzz target into the language understanding artificial intelligence (AI); - Generating application programming interface (API) calls and their arguments from the documentation by the language understanding AI; - Generating at least one fuzz driver from the API calls and their arguments.
[0061] This first step can be repeated and can be executed in synchronization with the following steps. Alternatively, the generation of the fuzz drivers is not executed synchronously, and the generated fuzz drivers are stored, for example, in a corpus. The corpus can function as a basis for the following steps.
[0062] In a second step, a fuzz driver is selected and executed (42) according to claim 7 of the training method 10, which includes the following (sub)steps. - A step of selecting at least one fuzz driver according to a selection strategy, the selection strategy including at least one of a first-in-first-out method derived by static analysis, the highest fitness value, and the deepest abstract syntax tree (AST), - A step of executing the selected fuzz driver against the fuzz target until an end criterion is met, the end criterion including at least one of no new coverage gain over several minutes, a crash being detected, and covering a code position not reached by the previous fuzz driver.
[0063] In a third step, a fitness value of the fuzz driver is generated (43) when the fuzz driver is executed. Optionally, the fitness value of the fuzz driver may be generated by measuring the generated code coverage of the executed fuzz driver. The fitness value is a measure of the success of fuzzing the fuzz target by the fuzz driver.
[0064] In a fourth step, the fitness value is fed back to the language understanding AI (44). The fitness value can be regarded as a reward for the training of the language understanding AI. In the fifth step, update the language understanding AI using the fitness value (45). For example, the weights of the AI or language model are updated with the fitness value (the value of the reward). As a result of this method, an AI or language model is obtained that is better trained, i.e., more reliable, with new unlabeled data. The unlabeled data may here be, for example, engine control unit code and optionally documentation from the corresponding source code.
[0065] According to one embodiment, the method further includes approximating the reward by executing only one of the tests in the automated test method. Thereby, the training can be accelerated.
Description of the reference signs
[0066] 21 Documentation 22 Language understanding artificial intelligence 23 Output 24 Fuzzy driver generator 25 Fuzzy driver 26 Fuzzy set-up 27 Fuzzy target 28 Feedback 29 Source code
Claims
1. A method (10) for generating a fuzz driver (25) for a fuzz setup (26), comprising the steps of: A step (11) of inputting a documentation (21) of a fuzz target (27) into a language understanding artificial intelligence (AI) (22); generating (12) application programming interface (API) calls and their arguments from said documentation (21) by said language understanding AI (22); generating (13) at least one fuzz driver (25) from said API calls and their arguments; A method (10), comprising:
2. The method of claim 1 , wherein the at least one generated fuzz driver (25) is input into a collection and / or corpus of fuzz drivers.
3. 3. The method of claim 1 or 2, wherein the documentation (21) comprises at least one of a handbook, a scope statement, a product requirements document, code, debug symbols of binaries, logs such as records of communication to and from a fuzz target, and / or programming comments.
4. 4. The method of claim 1, wherein the language understanding AI (22) comprises at least one of a large-scale language model (LLM), a natural language processing (NLP), a natural language understanding (NLU), and / or the language understanding AI is a trained model or a foundation model.
5. The method of claim 1 , further comprising, when generating API calls and their arguments, generating a fitness value for the API calls.
6. The method according to any one of claims 1 to 5, wherein the source code of the fuzz target (27) is input to the language understanding AI (22).
7. selecting at least one fuzz driver (25) according to a selection strategy, said selection strategy being heuristic or metric and / or including at least one of first in first out, first in last out, highest fitness value, deepest Abstract Syntax Tree (AST) derived by static analysis; running the selected fuzz driver (25) against the fuzz target (27) until a termination criterion is met, the termination criterion including at least one of no new coverage gains for several minutes, a crash being detected, and covering a code location not reached by a previous fuzz driver; The method of claim 1 , further comprising:
8. 8. The method of claim 7, wherein a fitness value of the fuzz driver is measured by a generated code coverage of the executed fuzz driver (25), and the fitness value is fed back to the language understanding AI (22), and / or the fitness value of the fuzz driver (25) is measured by a generated code coverage of the executed fuzz driver (25), and the fitness value is fed back to the generation of at least one fuzz driver (25).
9. A method for training a language understanding AI (22) to generate a fuzz driver (25) for a fuzz setup (26), comprising: Generating a fuzz driver (25) for a fuzz setup (26) according to claim 1; Selecting and executing a fuzz driver (25) according to claim 7; generating a fitness value for the fuzz driver (25) upon execution of the fuzz driver (25); feeding back said fitness value to said language understanding AI (22); updating the language understanding AI using the fitness value; A method comprising:
10. The method of claim 9 , wherein the fitness value of the fuzz driver (25) is generated by measuring the generated code coverage of the executed fuzz driver (25).
11. A computer system configured to carry out the method of any one of claims 1 to 10.
12. A computer program arranged to carry out the method according to any one of claims 1 to 10.
13. A computer readable medium or signal storing and / or including a computer program according to claim 12.