Fuzzy testing methods, apparatus and electronic equipment
By decompiling and analyzing closed-source binary programs, test drivers and samples that conform to the logic of binary programs are generated, solving the problem that fuzz testing in existing technologies depends on source code, and realizing efficient and accurate fuzz testing of closed-source binary programs.
Patent Information
- Application Number
- CN202610006876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-26
AI Technical Summary
In closed-source binary program scenarios, existing fuzzing methods rely on source code, which makes it difficult for test drivers to accurately understand binary semantics, reducing the effectiveness and automation of testing.
By decompiling closed-source binary programs, semantic information is extracted and encoded into target messages of the context protocol. The analysis is then performed using an intelligent agent to generate test drivers and test samples, including function metadata and data parsing mechanisms, ensuring that the generated test samples conform to the logic and format of the binary programs.
It enables accurate interpretation and fuzz testing of closed-source binary programs without relying on source code, improving testing efficiency and accuracy, generating high-quality test drivers and test samples, and enhancing the automation level of fuzz testing.
Smart Images

Figure CN122086759A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software security testing technology, and more specifically, to a method, apparatus, and electronic device for fuzz testing. Background Technology
[0002] Fuzz testing is an important software security testing technique that aims to detect potential security vulnerabilities by inputting large amounts of random or mutated data into software or systems. Related techniques utilize Large Language Models (LLMs) and source code knowledge graphs to enhance the input generation capabilities and semantic understanding of function calls, improving the targeting and effectiveness of fuzz testing. However, its limitation lies in its high dependence on the existence of source code, severely restricting its applicability to purely closed-source binary programs.
[0003] When faced with software scenarios without source code, the lack of key information in the source code makes it impossible for LLM to accurately understand function signatures, calling conventions, and parameter constraints. This can lead to the generated test drivers and test samples not matching the actual needs of the program, thereby reducing the effectiveness and automation of testing.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for fuzz testing, which at least solves the technical problem that the testing methods used in related technologies rely on the source code of binary programs to generate test drivers, which cannot correctly understand binary semantics in the context of closed-source binary programs, resulting in inaccurate test drivers.
[0006] According to one aspect of the embodiments of this application, a fuzzing method is provided, comprising: decompiling a target program to obtain semantic information of the target program, wherein the target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program; encoding the semantic information into a target message conforming to a context protocol; analyzing the target message using an intelligent agent to obtain a test driver and a test sample; and performing fuzzing on the target program using the test driver and the test sample.
[0007] In some embodiments of this application, decompiling the target program to obtain semantic information of the target program includes: decompiling the target program to obtain first semantic information and second semantic information, wherein the first semantic information includes function metadata for guiding the construction of the test driver, and the second semantic information includes a data parsing mechanism for guiding the generation of test samples; and using the first semantic information and the second semantic information as semantic information.
[0008] In some embodiments of this application, encoding semantic information into a target message conforming to a context protocol includes: encoding the semantic information using a preset format to obtain a target message, wherein the target message includes a Model Context Protocol (MCP) message, and the MCP message is used to explicitly convey the binary semantics of the target program.
[0009] In some embodiments of this application, the intelligent agent includes a first intelligent agent and a second intelligent agent; the analysis of the target message by the intelligent agent to obtain a test driver and a test sample includes: processing the target message and a first preset prompt word by the first intelligent agent to obtain a test driver, wherein the first preset prompt word is used to guide the first intelligent agent to generate executable code of the test driver corresponding to the target message; and processing the target message and the second preset prompt word by the second intelligent agent to obtain a test sample, wherein the second preset prompt word is used to guide the second intelligent agent to generate test sample metadata corresponding to the target message.
[0010] In some embodiments of this application, a second intelligent agent is used to process the target message and the second preset prompt word to obtain a test sample. This includes: using the target message and the second preset prompt word as input to the second intelligent agent to obtain test sample metadata; comparing the file format in the test sample metadata with a preset file format; if the file format matches the preset file format, driving the second intelligent agent to select a preset rule corresponding to the preset file format from a preset structured rule base to generate a test sample; if the file format does not match the preset file format, driving the second intelligent agent to generate a test sample corresponding to the binary semantics of the target message.
[0011] In some embodiments of this application, the method further includes: extracting platform information of the target program from the target message; determining a container image that matches the platform information, wherein the container image includes images corresponding to multiple systems respectively; configuring a test container corresponding to the container image, wherein the test container is used to compile a test driver.
[0012] In some embodiments of this application, after configuring a test container corresponding to the container image, the method further includes: compiling a test driver in the test container to obtain a compiled target test driver; setting a marker on the target data field of the test sample to obtain a target test sample, wherein the marker is used to verify whether the target function of the target program is called correctly and whether the parameters passed by the test driver to the target function flow through the expected path; injecting the target test sample into the input parameters of the target test driver, running the target test driver with the injected input parameters, and obtaining a verification result, wherein the verification result is used to indicate the validity of the test driver.
[0013] In some embodiments of this application, after obtaining the verification result, the method further includes: if the verification result indicates that the verification failed, regenerating the test driver corresponding to the target program according to the error type in the verification result; if the verification result indicates that the verification was successful, performing fuzz testing on the target program using the test driver and test samples.
[0014] According to another aspect of the embodiments of this application, a fuzzing apparatus is also provided, comprising: a decompilation module for decompiling a target program to obtain semantic information of the target program, wherein the target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program; an encoding module for encoding the semantic information into a target message conforming to a context protocol; an analysis module for analyzing the target message using an intelligent agent to obtain a test driver and test samples; and a testing module for performing fuzzing tests on the target program using the test driver and test samples.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the method for implementing the above-described fuzz test.
[0016] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-described fuzzing test method by running the computer program.
[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the above-described fuzzing method.
[0018] In this embodiment, by converting the decompiled semantic information of a closed-source binary program into a standardized, context-sensitive message format, and using an intelligent agent with advanced reasoning and code generation capabilities to parse this semantic information, targeted test drivers and structured test samples are automatically generated for fuzz testing. This achieves the goal of accurately interpreting binary semantics to generate accurate test drivers, thereby improving the efficiency and accuracy of fuzz testing on pure closed-source binary programs. It also solves the technical problem that the test methods used in related technologies rely on the source code of the binary program to generate test drivers, which cannot correctly understand binary semantics in the context of closed-source binary programs, resulting in inaccurate test drivers. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a hardware structure block diagram of a computer terminal for a fuzz testing method according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of a fuzz testing method according to an embodiment of this application;
[0022] Figure 3 This is a system architecture diagram of a fuzzing method according to an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of a fuzz testing apparatus according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0027] Model Context Protocol (MCP) is a standardized communication protocol used to accurately express and transmit the semantic information of closed-source binary programs extracted by decompilation tools to the agent, ensuring that the agent can accurately understand the program's logical structure, parameter meanings, calling conventions, and potential security sensitivities. In this embodiment, MCP serves as a crucial bridge connecting the decompiled output and the agent, enabling the agent to generate effective test drivers and test samples based on accurate contextual understanding, thereby significantly improving the efficiency and accuracy of fuzz testing.
[0028] Test Harness (or simply harness): In the field of software testing, a test harness specifically refers to a set of programs or code created to test specific functions or components. These programs simulate external input to elicit the behavior of the target software, allowing for observation and analysis of its responses. In this embodiment, the test harness is automatically generated by an agent based on MCP messages and is specifically designed to load and execute inputs from closed-source binary programs, ensuring that tests cover critical parts of the target code without being limited by the availability of the source code.
[0029] Structured test cases (or simply test cases): Unlike random or unstructured inputs, structured test cases are carefully designed test inputs that follow a specific format or protocol. They aim to maximize the triggering of complex logical paths in the target program, thereby discovering potential vulnerabilities or abnormal behaviors. In this embodiment, the agent generates structured test cases based on the program semantics contained in the MCP. These samples can more effectively bypass the initial screening stage and directly reach the program's internal workings, significantly increasing the probability of vulnerability triggering during fuzzing.
[0030] The agent is an intelligent proxy that integrates large-scale language models, automated tool invocation, and advanced task planning capabilities. It is capable of performing complex code understanding and generation tasks in fuzzing scenarios. In this embodiment, the agent plays a core role. It receives semantic information in MCP format, and through multi-step reasoning and code generation, produces test drivers and structured test samples. This automates the fuzzing of closed-source binary programs, reducing the need for manual intervention and enhancing the comprehensiveness and accuracy of the tests. It should be noted that the agent in this application is not limited to the agent that produces test drivers and structured test samples.
[0031] Dynamic Taint Analysis (VTA) is a security analysis technique that identifies how data is processed and whether potential security vulnerabilities exist by tracing the propagation path of specific data values during program execution. In this embodiment, VTA is used to verify whether the generated test driver correctly calls the target function of the closed-source binary program and ensures that the test input data flows along the expected path, thereby ensuring the effectiveness of fuzz testing. It also provides real-time feedback for the learning and optimization of the agent.
[0032] Containerized deployment: A software deployment method based on container technology (such as Docker) that packages an application and its dependencies into a lightweight, portable container, ensuring consistent operation in any environment. In this embodiment, containerized deployment is used to build isolated test environments that automatically adapt to the target program's runtime platform (Windows, Linux, or macOS), thereby eliminating compilation environment compatibility issues and ensuring that test drivers and binary target programs execute in the same or similar environments, enhancing the reliability and convenience of fuzz testing.
[0033] In the field of fuzzing technology, especially for fuzzing closed-source binary programs, there are many technical challenges and limitations. These limitations hinder the widespread application and efficiency improvement of fuzzing in the field of software security, including:
[0034] (1) Traditional gray-box fuzzers (such as AFL-QEMU): AFL-QEMU is a combination of American Fuzzy Lop (AFL) and the QEMU user-space emulator. AFL is a gray-box fuzzer based on coverage feedback, which records the program execution path (basic block transition) through instrumentation and uses it to guide the mutation strategy; QEMU is an open-source CPU emulator that supports cross-architecture execution (such as running ARM programs on x86 hosts). In the AFL-QEMU mode, AFL utilizes QEMU's dynamic binary translation (DBT) capability to transparently insert lightweight instrumentation code without modifying the target binary, thereby achieving coverage feedback for closed-source programs.
[0035] Its advantages include the absence of source code, support for arbitrary closed-source binaries, and direct fuzzing of .exe (Windows PE), .so (Linux ELF), firmware, embedded programs, etc., without recompilation or instrumentation, greatly lowering the barrier to entry. However, its shortcomings, such as lack of semantic understanding, high performance overhead, and limitation to crash detection, coupled with the need for manual development of harnesses, have prompted researchers to explore more intelligent alternatives (such as protocol-aware fuzzing and LLM-assisted fuzzing).
[0036] (2) Fuzz4All: A solution for general fuzz testing using large language models, supporting multiple programming languages (such as C / C++, Java, Python, Go, SMT2) and related systems. Fuzz4All introduces autoprompting technology, which can automatically extract potentially lengthy inputs provided by users (such as technical documents, sample code, user manuals, etc.) into prompts suitable for fuzz testing. Thanks to pre-training on massive amounts of code data, LLM can generate diverse test inputs that conform to the grammatical and semantic constraints of the language.
[0037] However, its reliance on source code, technical documents, manuals, etc. to extract fuzzy test hints or build knowledge graphs, coupled with the need for manual development and writing of harnesses, limits its application in purely closed-source binary scenarios. In addition, the inherent limitations of LLM and the illusion problem inherent in large language models may lead to generated code containing syntactic or semantic errors. Although there are dynamic repair mechanisms, invalid test drivers or mutations may still be generated, affecting test efficiency.
[0038] (3) CKGFuzzer: Knowledge graph enhancement driven generation. Through code knowledge graph, LLM can better understand the calling relationship and constraints between APIs, generate semantically valid test code, and can self-repair and optimize. It has dynamic program repair capabilities, can automatically correct the syntax and API usage errors generated by LLM, and optimize tests using coverage feedback. However, it depends on the source code and requires the source code of the target software to build the knowledge graph, which is limited in pure closed-source binary scenarios.
[0039] (4) AutoHarness: In traditional fuzzing, 90% of closed-source programs cannot be fuzzed due to the lack of a harness. AutoHarness can automatically analyze function signatures, parameter types, and dependencies to generate compilable C / C++ driver code, significantly lowering the barrier to entry. The generated harness can correctly call the target function, enabling the fuzzer to truly test the deep logic rather than remaining at the input parsing level. It can also be integrated with mainstream fuzzers, typically outputting harnesses compatible with LibFuzzer or AFL, seamlessly integrating into the existing fuzzing ecosystem. However, it does not support closed-source binary programs; currently, source code is generally required to analyze function signatures, parameter types, and dependencies to generate usable driver code harnesses.
[0040] In summary, although large model (LLM) has been used for code generation, the lack of a standardized program semantic input protocol leads to inaccurate understanding of binary semantics, and the generated harness often fails due to type errors and calling convention mismatches.
[0041] To address the aforementioned technical problems, this application provides corresponding solutions, which are detailed below.
[0042] The fuzz testing method embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal for implementing a fuzz testing method is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions connected via wired and / or wireless networks. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0043] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fuzzing method in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned fuzzing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The transmission module 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 106 may be a radio frequency (RF) module, used for wireless communication with the Internet.
[0046] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.
[0047] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer terminal shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computer terminal.
[0048] In the above operating environment, this application provides a method embodiment for fuzz testing. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0049] Figure 2 This is a flowchart of a fuzz testing method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0050] Step S202: Decompile the target program to obtain the semantic information of the target program. The target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program.
[0051] In step S202 above, decompilation is the process of converting machine code or bytecode back into a high-level language representation that is close to the original source code. This is used to extract understandable semantic information from closed-source binary programs. Semantic information is data that describes the program's functions, parameter meanings, data flow, control flow, and other internal logic. It is extracted from closed-source binary programs through decompilation and used by intelligent agents to understand the program's data processing logic.
[0052] It should be noted that closed-source binary programs refer to software programs that are distributed and executed only in binary form without providing source code, and are the main test objects of this application's embodiments.
[0053] In some embodiments of this application, professional decompilation tools such as Ghidra / IDA can be used to process closed-source binary programs and output readable pseudocode or intermediate representation (IR). Key semantic information, including function signatures, parameter types, calling conventions, string references, and dangerous functions called, can be extracted from the decompiled output and converted into a preset format.
[0054] To address the issue of inaccurate understanding of binary semantics by intelligent agents, the target program can be decompiled to obtain its semantic information in the following way: the target program is decompiled to obtain first semantic information and second semantic information, where the first semantic information includes function metadata used to guide the construction of the test driver, and the second semantic information includes a data parsing mechanism used to guide the generation of test samples; the first and second semantic information are then used as semantic information.
[0055] Specifically, the first semantic information refers to the information about function metadata extracted from the decompiled results, including but not limited to function name, parameter list (type, order, default value), return type, etc., used to guide the agent in building test drivers and ensure that the generated harness can correctly call the target function and meet its type and calling convention requirements. The second semantic information refers to the data parsing mechanism information collected from the decompiled code, involving how to process and parse specific types of input data (such as file formats, network protocols, etc.), used to guide the agent in generating structured test cases, ensuring that the generated test cases are in the correct format and can effectively trigger the data processing logic of the target binary program.
[0056] In the above process, the first semantic information provides the agent with explicit function metadata, enabling it to generate compliant test drivers and reduce type errors and inconsistencies in calling conventions. The second semantic information provides a deep understanding of the input data processing mechanism, allowing the agent to generate highly structured and correctly formatted test samples, significantly improving the coverage of the deep logic of the target binary program by fuzzing. Through standardized communication protocols, the accuracy of the agent's output is significantly improved, achieving the goal of generating high-quality test drivers and test samples without relying on source code, thereby enhancing the efficiency and effectiveness of fuzzing closed-source programs.
[0057] It should be noted that semantic information can include not only the first semantic information and the second semantic information mentioned above, but also overlapping parts of the first semantic information and the second semantic information, which is not limited here.
[0058] To facilitate understanding of the above decompilation process, some specific examples will be used for explanation below.
[0059] Specifically, a binary program analysis agent (used for decompilation and MCP construction) can be used, and Ghidra / IDA can be used to decompile closed-source binaries to extract system platform information, assembly code, decompiled IR structured output, pseudo-C code, etc. of the target program. Among these, the following information is of particular interest:
[0060] a. Function address and name (or symbol);
[0061] b. Number and type of parameters (e.g., char) (size_t), direction (in / out);
[0062] c. String references (such as "PNG", "system");
[0063] d. Calling dangerous functions (such as memcpy, system);
[0064] e. Infer the file / protocol type to be processed (e.g., "PNG", "IDAT", "HTTP / 1.1");
[0065] f. Internally called parsing functions (such as "parse_png_chunk", "json_parse");
[0066] g. Magic number, version number, etc. (e.g., 0x89504E47).
[0067] The above information can then be encoded into a target message (such as an MCP message) according to a predefined schema.
[0068] Step S204: Encode the semantic information into a target message that conforms to the context protocol.
[0069] In step S204 above, the context protocol is a customized protocol used to encode the decompiled results of the binary program in a structured, typed, and semantic way so that the agent can understand and process them. The target message refers to the encoded semantic information. This message format can not only contain function metadata and data parsing mechanisms, but also cover high-level information that is crucial to generating test drivers and samples, such as strings, magic numbers, and file types.
[0070] Since the decompiled output contains a large number of uncertain types (such as undefined4), directly allowing the agent to process them will lead to a high error rate. However, by using context protocol encoding, the agent can obtain a clear context, including explicit parameter types in the function signature, calling conventions, parameter semantics, and security-related clues. This improves its ability to generate test drivers and samples and reduces invalid tests caused by type errors and calling convention mismatches.
[0071] To structure and type the information obtained from decompilation, semantic information can be encoded into target messages conforming to the context protocol in the following way: the semantic information is encoded using a preset format to obtain target messages, where the target messages include Model Context Protocol (MCP) messages, which are used to explicitly convey the binary semantics of the target program.
[0072] Model Context Protocol (MCP) is a protocol for communication between agents and program analysis tools. It allows agents to receive and understand complex program logic and semantic details, especially for closed-source binary programs. Specifically, an MCP encoder can be developed to understand the pseudocode or IR output by decompilers and convert it into MCP messages. The MCP message format predefines a series of tags and fields, such as function names, parameter lists, type annotations, strings, and magic number tags, ensuring that the information received by the agent has clear type and contextual relevance. Alternatively, an MCP schema can be defined to structure and type the information obtained from decompilation, ensuring that the information received by the AI agent is organized and parsable. The schema contains data points such as function addresses, parameter types and semantics, and known strings or magic numbers. This information is encoded as part of the MCP message, facilitating the agent's understanding and utilization. Through the MCP schema, the problems of "type confusion" and "semantic ambiguity" encountered by traditional LLMs when processing binary programs can be solved, making test-driven and sample generation more accurate.
[0073] By encoding semantic information into MCP messages, the key problem of inaccurate understanding of binary semantics by agents in closed-source binary program testing is solved. The introduction of MCP messages significantly enhances the accuracy of agent-generated harnesses by explicitly passing information such as function signatures, parameter types, calling conventions, and security clues (such as strings and magic numbers). This is because it provides sufficient context and semantic clues, enabling agents to understand and simulate the internal behavior of closed-source binary programs in greater detail.
[0074] Step S206: The intelligent agent analyzes the target message to obtain the test driver and test sample.
[0075] In step S206 above, the intelligent agent includes an AI agent, which combines the language understanding ability of a large model with the execution ability of automated tools, and is able to complete high-level tasks such as code generation, driver construction, and test sample creation.
[0076] A test driver is a specific code segment used to load and execute a target program in a test environment. Its main function is to provide the input and context environment required for the program to run normally, ensuring that the target function can be called correctly for testing. A test sample is a set of input data used to verify the functionality and robustness of a software system or module (such as a target program).
[0077] In some embodiments of this application, after receiving the MCP message, the agent performs multi-level analysis and understanding: First, the agent parses the function metadata in the MCP, including function name, parameter type, return type, etc., to construct the skeleton of the test driver; then, the agent generates specific code based on parameter semantics and calling conventions to ensure that the generated test driver can accurately call the target function.
[0078] Furthermore, the agent generates structured test samples based on security clues and data parsing mechanism information in the MCP message. For example, if the MCP message indicates that the objective function is a PNG image parser, the agent will generate an image file containing the correct PNG magic number and format as the initial template for the test sample. Then, the agent will further refine and mutate this template based on the data format that the target program is expected to process, generating a series of test cases for fuzz testing.
[0079] The above steps, by utilizing AI agents to analyze and process MCP messages, overcome the limitations of relying on source code to generate test drivers in traditional closed-source binary program fuzzing. Specifically, the agent can accurately understand the functional requirements and internal logic of the target program through a pre-defined MCP message structure, thereby generating highly accurate test drivers and test samples.
[0080] To address the challenge of generating test drivers and test samples in closed-source binary program scenarios, test drivers and test samples can be generated as follows: The intelligent agents include a first intelligent agent and a second intelligent agent. Based on this, the first intelligent agent processes the target message and a first preset prompt to obtain the test driver, where the first preset prompt guides the first intelligent agent to generate executable code for the test driver corresponding to the target message. The second intelligent agent processes the target message and the second preset prompt to obtain the test sample, where the second preset prompt guides the second intelligent agent to generate test sample metadata corresponding to the target message.
[0081] The first agent is primarily responsible for analyzing and generating test driver code from the target message. It is designed to understand and execute specific types of prompts to ensure that the generated test driver accurately reflects the binary semantics of the target program and meets testing requirements. The first preset prompt contains specific instructions and semantics to help the agent understand how to generate correct executable code based on the function metadata, calling conventions, and other information in the target message.
[0082] Specifically, the first agent receives a target message encoded by MCP and a first preset prompt word. Based on the prompt word and function metadata in the target message, such as function name, parameter types, and calling conventions, the agent generates compilable test driver code. This process solves the problem in traditional methods where fuzzing of closed-source binary programs lacks accurate function signatures and calling conventions, ensuring that the generated test driver can correctly and error-free call the target function.
[0083] Unlike the first agent, the second agent focuses on analyzing the target message and generating metadata for test samples. This metadata includes the format, structure, and key data points of the test samples, providing guidance for generating specific test samples later. By parsing the high-level semantic information in the MCP, the second agent can intelligently identify the data types and formats processed by the target program, thereby generating more effective test samples. Similar to the first preset prompts, the second preset prompts guide the second agent in generating metadata for test samples. For example, it may include information on how to infer the input format type processed by the target function based on the MCP message.
[0084] Specifically, after receiving the target message and the second preset prompt, the second agent, based on the security clues and data parsing mechanism in the target message and combined with the generation strategy in the prompt, derives the metadata of the test sample. For example, it may include the initial structure of the test sample (such as the basic format of a PNG image), key fields (such as the width and height of the image), and mutation points (such as modifying the magic number or adding a special data stream). This process solves the problems of low efficiency and blindness in test sample generation in traditional fuzzing due to the lack of structured input guidance. Through precise metadata description, the effectiveness of the test sample and the probability of vulnerability triggering are significantly improved.
[0085] By combining the first intelligent agent and the first preset prompt, the AI can generate compliant test driver code based on the function metadata and calling conventions obtained from decompilation analysis. Similarly, the cooperation of the second intelligent agent and the second preset prompt enables the intelligent agent to intelligently generate the metadata of the test sample based on the high-level semantic information extracted from the MCP, such as file or protocol type, key strings and magic number. Through predefined rules and mutation strategies, the second intelligent agent can ensure that the test sample is not only correctly formatted, but also can trigger the key logic path of the target program.
[0086] To generate test samples that match the target program more efficiently and accurately, test samples can be determined as follows: The target message and a second preset prompt are used as input to the second agent to obtain test sample metadata; the file format in the test sample metadata is compared with a preset file format; if the file format matches the preset file format, the second agent is driven to select a preset rule corresponding to the preset file format from a pre-set structured rule base to generate a test sample; if the file format does not match the preset file format, the second agent is driven to generate a test sample corresponding to the binary semantics of the target message.
[0087] It's important to note that the test sample metadata contains all the crucial information needed to generate test samples, such as file format metadata, data structure descriptions, value ranges for specific fields, and necessary header information. This metadata forms the basis for the second agent's test sample generation, determining the format and content of the test samples. Preset file formats are a collection of file formats commonly encountered in fuzzing, such as PNG, JPEG, and PDF. Each preset file format has its specific rules and structure. The existence of preset file formats ensures that the second agent adheres to correct file format specifications when generating test samples, avoiding the generation of invalid or incompatible test samples. The pre-defined structured rule base is a database used to store various preset file format rules. The rule base includes detailed descriptions of file formats, validity verification rules, and the structure and range of key fields. The second agent consults this rule base when generating test samples to ensure that the generated test samples meet the expected file format requirements.
[0088] Specifically, after receiving the target message and the second preset prompt, the second agent first parses the MCP message, extracting semantic information related to file format, protocol type, key parameters, etc., to construct the metadata of the test sample. The second agent compares the file format in the constructed test sample metadata with the preset file format. If they match, the agent generates a test sample according to the corresponding rules in the preset structured rule base; if they do not match, the agent automatically generates a test sample based on the binary semantic information in the target message.
[0089] The above approach addresses the lack of guidance and structure in test sample generation during fuzzing. With the assistance of preset file formats and rule bases, the agent can generate test samples that match the target program more efficiently and accurately. It also handles cases where the preset format cannot be directly matched, maintaining the flexibility and comprehensiveness of the test. Specifically, the introduction of preset file formats and rule bases increases the standardization and efficiency of test sample generation. That is, by comparing the test sample metadata with the preset file format, and generating test samples using a pre-defined structured rule base when a match is found, it ensures that these samples adhere to the file format specifications, reducing invalid tests and accelerating the generation of test samples, thus improving the automation of the entire fuzzing process. The adaptive test sample generation mechanism further enhances the flexibility of fuzzing. For cases where the preset file format cannot be directly matched, the second agent can automatically generate test samples based on the binary semantic information in the target message. This adaptive mechanism not only broadens the applicability of fuzzing but also generates customized test samples for specific closed-source binary programs, further improving the depth and breadth of the test.
[0090] To facilitate understanding of the construction process of the test driver and test samples described above, the following explanation will be provided in conjunction with some specific examples.
[0091] Specifically, a test-driven agent can be used to generate an agent that uses MCP messages as context input to the AI agent (such as a fine-tuned CodeLlama or GPT-4), along with a structured prompt (i.e., the first preset prompt word) as follows:
[0092] You are a Windows binary expert. Please generate a LibFuzzer-compatible harness for the following MCP:
[0093] {……MCP content……}.
[0094] Requirements: Use Windows call to read; initialize necessary global state; read input from stdin as buf.
[0095] By constructing the above prompts, the AI can output a valid harness code.
[0096] It should be noted that the aforementioned first preset prompt words, by specifying the platform (Windows) and the format and functional standards of the harness (LibFuzzer compatible), ensure that the generated harness can be seamlessly integrated with LibFuzzer, thus allowing it to be directly used in the fuzzing process without additional format conversion or adaptation. Furthermore, specific technical guidance and environment settings for generating the harness are provided through specific quality requirements.
[0097] For generating test samples, a sample generation agent can be used. The AI agent utilizes high-level semantic information in the MCP (Multi-Step Prompt) as context to perform multi-step reasoning and generate structured samples that conform to the expected format of the target program, thereby significantly improving the efficiency of vulnerability triggering. The second preset prompt word could be, for example:
[0098] You are a binary security expert. Based on the following MCP description, infer the input format type processed by the objective function:
[0099] {...MCP content... (such as function, string, calls, params, etc.)}.
[0100] Please answer in the following format: {"format": "PNG", "confidence": "high"}.
[0101] Based on the LLM output, the AI Agent instantiates a legitimate sample. If the file type matches the pre-set structured rule base, it generates a sample according to the rules; otherwise, it generates a sample based on the program semantics and uses it as an initial seed sample for fuzz testing.
[0102] It should be noted that the requirement for the "confidence" field in the second preset prompt lays the foundation for establishing a closed-loop feedback mechanism. Based on the level of "confidence," it can be determined whether to accept the test samples generated by the AI Agent, or to adjust the MCP content or the inference strategy of the intelligent agent, thereby continuously optimizing the quality of the test samples.
[0103] In the case of no match, for example, the initial seed sample can be generated through the following steps:
[0104] a) Fill magic number: 89 50 4E 47 ...;
[0105] b) Construct IHDR: width=100, height=100, ...;
[0106] c) Add IDAT: pixel data compressed with valid zlib;
[0107] d) Calculate CRC;
[0108] e) Output: A valid seed.png file (binary).
[0109] Specifically, the magic number is a specific sequence of bytes used to identify the type or format of a file. Magic number information contained in the MCP message, such as "89 50 4E 47...", is the magic number for a PNG image file. When generating test samples, the agent first fills in these magic numbers to ensure the file header conforms to the target program's expected format. For image files like PNGs, the IHDR is a segment storing basic image information (such as width, height, color type, bit depth, etc.). The agent needs to construct a reasonable IHDR based on the information revealed in the MCP, such as the image size or format preferences that the objective function may involve. For example, it might set a segment with a width and height of 100, while ensuring that other parameters (such as bit depth and color type) also conform to the target program's expectations.
[0110] The IDAT segment contains the actual pixel data of the image. For PNG files, this data is typically compressed using zlib. The agent fills the IDAT segment by generating appropriate, valid zlib-compressed data, such as using the pixel values of a simple image or pattern, processed by the zlib compression algorithm, and then filling it in. CRC is a commonly used checksum used to detect errors in data transmission. In many file formats, especially those concerned with data integrity, such as PNG, CRC is used to verify the integrity of each segment of data. After generating test samples, the agent needs to calculate and add the correct CRC value to the end of the corresponding segment to ensure the integrity and validity of the file. The CRC calculation is based on the data within the segment; therefore, the agent needs to recalculate the CRC every time data is generated or mutated. This is a necessary condition to ensure that the test samples can be correctly parsed and processed in the target program.
[0111] After completing the filling and calculation of the header, structure, data, and CRC, the agent combines all this information into a complete and valid PNG file (or other equivalent format file) and saves it as a binary seed.png file. This serves as the initial seed sample for fuzzing and is the starting point of the testing process. The valid seed sample generated by the agent not only meets the target program's basic requirements for input format but also contains elements that could potentially trigger vulnerabilities based on program semantics, providing a highly targeted starting point for subsequent fuzzing.
[0112] In the case of a match, file format rules can be loaded. The system has a pre-built structured rule library containing common file formats (such as png, jpeg, pdf, doc, xls, etc.). The AI Agent loads the corresponding rules based on the inferred format (such as PNG) and generates a precise mutation strategy for the fuzzer to use.
[0113] Specifically, once the file format (such as PNG) is determined, the AI Agent loads the corresponding rules from the system's pre-built structured rule base. For example, for the PNG format, the rules may include how to construct segments such as IHDR and IDAT, as well as their expected attributes and value ranges. The existence of the pre-built rule base not only provides authoritative information on the file format, but also simplifies the AI Agent's generation task, allowing it to focus on formulating mutation strategies rather than building the file structure from scratch.
[0114] Based on the loaded file format rules, the AI Agent analyzes the mutable points defined in the rules, such as the width, height, color type, and compression level of a PNG image. These mutable points are key locations for mutating inputs in fuzz testing. By controlling the changes in these points, more execution paths can be covered, increasing the probability of discovering potential vulnerabilities. Based on the analysis results of the mutable points, combined with information about the objective function provided in the MCP (such as parameter types and calling conventions), the AI Agent formulates a precise mutation strategy. This strategy may include mutation frequency, mutation range, and mutation methods. For example, for the width and height of a PNG image, a smaller mutation range, such as ±10%, can be set to maintain the basic structure of the image; while for the color type or compression level, more mutation options can be tried to test the program behavior under different conditions.
[0115] Ultimately, based on the mutation strategy, the AI Agent generates a series of test samples. Each sample is constructed under the guidance of a pre-defined rule base and finely adjusted according to the mutation strategy to ensure sample diversity. For example, for PNG format, the AI Agent may generate multiple images, each slightly different in dimensions such as width, height, and color type, but all conforming to the PNG specification. These diverse samples will serve as seeds for fuzz testing, further mutating and testing to discover security vulnerabilities in the program.
[0116] Step S208: Perform fuzz testing on the target program using a test driver and test samples.
[0117] In step S208 above, during the fuzzing preparation phase, the test driver is compiled and configured to meet the specific requirements of the target binary program. For example, if the target program is a Windows PE format binary, the test driver will be built as C / C++ code that can run in the corresponding Windows Docker image. After the test driver is correctly loaded, the fuzzer (such as LibFuzzer or AFL) uses the test cases generated by the AI Agent as input to repeatedly mutate and test the target program.
[0118] To ensure that the test driver can be compiled and run correctly, the following steps can be performed: extract the platform information of the target program from the target message; determine the container image that matches the platform information, wherein the container image includes images corresponding to multiple systems respectively; configure the test container corresponding to the container image, wherein the test container is used to compile the test driver.
[0119] It's important to note that platform information refers to the target program's operating system, architecture (e.g., x86, ARM), compiler, and runtime environment. A container image is a package containing all the dependencies required to run an application, including the operating system, libraries, tools, and configuration. In fuzzing, container images provide an isolated, controlled environment for compiling and running test drivers and target programs for fuzzing. Different operating systems and architectures have corresponding container images; for example, for Windows PE format binaries, a Windows operating system container image will be used. A test container is a running instance created based on a selected container image.
[0120] Specifically, the agent parses the MCP message to identify key platform-related information, such as the format of the target binary file (ELF, PE, Mach-O, etc.). This typically reflects the platform environment of the target program; for example, an ELF format binary file means the program runs on Linux, while a PE format file points to a Windows platform. Furthermore, the agent may infer more platform details from other indirect clues in the MCP, such as calling conventions (__cdecl, __stdcall, etc.), linker flags, and compiler version.
[0121] Based on the extracted platform information, the agent selects the most suitable image from a pre-configured container image library. For example, if the target message indicates that the program is based on Windows PE format, the agent will select a Docker image containing a Windows environment to ensure that the test driver can be compiled and run in the correct environment. Considering that fuzzing may require specific tools and libraries, the agent may further filter or customize the container image. For example, adding specific fuzzing tools (such as AFL, LibFuzzer), dynamic analysis tools (such as Pin, Valgrind), and related development libraries (such as OpenSSL, freetype, etc.) will improve the efficiency and accuracy of fuzzing.
[0122] Finally, when creating the test container, the agent will configure the OS type, architecture, necessary environment variables, and dependent libraries in the container image. For example, for a Windows PE format binary program, the test container will be based on the Windows image and configured with compilation tools (such as MSVC), linker, and runtime libraries to ensure that the test driver can be compiled successfully and run in the target container.
[0123] Through the above steps, the agent can accurately understand the platform characteristics of the target program through the MCP (Target Message) and automatically select or customize from the container image library based on the platform information. This avoids the tedious work of manually selecting and configuring container images, reduces the possibility of operational errors, and also provides a unified testing environment for closed-source binary programs of different formats, supporting diverse closed-source program testing needs such as Windows PE, Linux ELF, and macOS Mach-O.
[0124] In some embodiments of this application, after configuring a test container corresponding to the container image, the following steps can also be performed: compiling a test driver in the test container to obtain a compiled target test driver; setting a marker on the target data field of the test sample to obtain a target test sample, wherein the marker is used to verify whether the target function of the target program is called correctly and whether the parameters passed by the test driver to the target function flow through the expected path; injecting the target test sample into the input parameters of the target test driver, running the target test driver with the injected input parameters, and obtaining a verification result, wherein the verification result is used to indicate the validity of the test driver.
[0125] It should be noted that in the field of fuzz testing, tagging is a data tracking technique used to attach a flag to a specific field of a test sample. This flag can be identified and tracked by dynamic taint analysis tools. By setting a flag on the target data field, the flow of parameters within the target program can be monitored, verifying whether the test driver successfully guides the execution of the target function and whether the parameters are delivered according to the expected path.
[0126] Specifically, the test driver is written based on the function signature and calling convention described in the MCP message, and then compiled in the test container. Simultaneously, specific fields of the test sample, such as function parameters and important data structure members, are marked. These marks are treated as special data tags during dynamic analysis. For example, once the width and height parameters of a PNG image are marked, the dynamic analysis tool will track how these parameters are processed and used during the execution of the target program. The marked target test sample is injected into the input parameters of the target test driver, and then the test driver is run in the test container. Dynamic analysis tools (such as Intel Processor Trace, Pin, etc.) track the marked data flow in real time, recording the call details of the target function, the parameter flow path, and whether any unexpected abnormal behavior is triggered. The verification results directly reflect whether the test driver successfully guided the target program to its expected execution path.
[0127] The analysis of verification results is a closed-loop feedback process. If the target function is not called, or the path of parameter flow does not match the expectation, the agent can adjust the generation strategy of the test driver based on the verification results. For example, it can optimize the prompt (i.e., the first / second preset prompt words), improve the function call simulation, or enhance the semantic richness of the MCP message. Through repeated iterations, an effective test driver can be obtained that can accurately guide the execution of the target function, ensuring the effectiveness and depth of the test.
[0128] Specifically, after obtaining the verification results, the following steps can be performed: if the verification result indicates that the verification failed, regenerate the test driver corresponding to the target program according to the error type in the verification result; if the verification result indicates that the verification was successful, perform fuzz testing on the target program using the test driver and test samples.
[0129] In other words, when the verification result indicates that the verification has failed, the AI Agent will analyze the error type in the verification result (such as parameter type mismatch, calling convention error, incorrect data flow, etc.), and based on these specific errors, the AI Agent will adjust and optimize the prompt, MCP message, etc., and regenerate the test driver code. For example, if the verification fails because of parameter type identification error, the MCP message will contain more detailed type information in the next version to guide the AI Agent to generate the correct type conversion code and ensure that the parameters can be correctly passed to the target function.
[0130] If the verification result indicates successful verification, it means that the test driver has been able to correctly guide the execution of the target function and process the test samples. At this point, the test driver and test samples will be used in the formal fuzzing process. The agent no longer needs to modify the test driver, but instead turns to generating more test samples to cover more execution paths and potential vulnerabilities.
[0131] Through the above steps, the verification results are not only used to guide the generation of test drivers, but also to allow the agent to dynamically adjust the prompt based on the error type, forming a closed-loop feedback. This means that even if the initial test driver has deviations due to insufficient understanding of binary semantics, it can still generate an accurate test driver through multiple iterations and optimizations, greatly improving the efficiency and success rate of fuzz testing for closed-source programs. Furthermore, based on the information provided in the MCP message, the AI Agent can intelligently generate test drivers. Even after the initial verification fails, the agent can regenerate the test driver based on the specific error type, such as a mismatch in calling conventions or an anomaly in the data flow, ensuring that it is consistent with the calling mechanism and data flow direction of the target program. This solves the problem in traditional fuzzing where test driver generation depends on the source code.
[0132] To facilitate understanding of the fuzz testing process described above, the following explanation will be provided in conjunction with some specific examples.
[0133] Specifically, multiple agents can be used to execute different test-related actions, including:
[0134] (1) Deploy Agent in Test Environment: Based on the target program platform information identified in the decompilation analysis stage, deploy the test environment using containerization, select different system Docker containers (Windows / Linux / macOS) for subsequent compilation of the harness program, verification of the effectiveness of the test driver harness, execution of fuzz tests, etc.
[0135] (2) Validation Agent: In the deployed test Docker container, compile Harness and run it for the first time, inject AI-generated samples and mark them with taints (such as marking data as tainted), and verify them through lightweight dynamic analysis (Intel Processor Trace + custom tracer):
[0136] Whether ParseConfig is called; whether data is passed as the first parameter; whether the expected data stream exists (such as entering the closed-source binary program entry point).
[0137] If verification fails, the AI Agent adjusts the harness generation strategy and regenerates the harness program code based on the feedback error type (such as "parameter not passed"), repeating the compilation process until verification is passed.
[0138] (3) Fuzz Agent: Used to perform fuzz testing and continuously optimize. After verification, start a fuzz tester such as LibFuzzer / AFL++ and use AI-generated seeds (such as binary templates that conform to the protocol structure) to perform testing. When a new path is found or a crash occurs, the MCP can be updated in reverse (such as supplementing newly identified parameter constraints) to form a closed loop.
[0139] In addition, feedback-driven test optimization can be performed, which means using the edge coverage statistics module to analyze test execution paths, implement heat redirection strategies, intelligently mutate samples, conduct targeted testing on low-density execution paths, and optimize test case sets by dynamically adjusting test intensity parameters and environment-aware scheduling to improve vulnerability discovery efficiency.
[0140] To perform vulnerability verification and exploitability assessment, a vulnerability analysis agent can be introduced to verify the validity of the captured anomalous behavior using Proof-of-Concept (PoC). By comparing the anomalies between the simulation environment and the real hardware through differential testing, and combining CVSS scoring and attack tree models to quantify vulnerability risks, a detailed vulnerability report can be generated, providing a basis for remediation and defense.
[0141] Through steps S202 to S208, the decompiled semantic information of the closed-source binary program is transformed into a standardized, context-sensitive message format. An intelligent agent with advanced reasoning and code generation capabilities is used to parse this semantic information, automatically generating targeted test drivers and structured test samples for fuzz testing. This achieves the goal of accurately interpreting binary semantics to generate accurate test drivers, thereby improving the efficiency and accuracy of fuzz testing on pure closed-source binary programs. It also solves the technical problem that the test methods used in related technologies rely on the source code of the binary program to generate test drivers, which cannot correctly understand binary semantics in the context of closed-source binary programs, resulting in inaccurate test drivers.
[0142] Figure 3 This is a system architecture for a fuzz testing method according to an embodiment of this application, such as... Figure 3 As shown, the system includes:
[0143] Input: Closed-source binary programs (pe / elf / mach-o), that is, any closed-source program in PE (Windows PortableExecutable), ELF (Executable and Linkable Format, Linux and Unix-like systems), or Mach-O (Macintosh Operating System Object file format, macOS) format.
[0144] System processing includes:
[0145] 1. Binary program analysis agent, including: decompiling semantic extraction of IDA / Ghidra assembly / IR / pseudo-C code; MCPBuilder agent to build standardized MCP; identification of formats and platforms.
[0146] Specifically, decompilation tools (such as IDA Pro or Ghidra) can be used to extract assembly code, intermediate representation (IR), pseudo-C code, and other information from closed-source binary programs, providing a semantic foundation for subsequent MCP construction and test driver generation. The MCPBuilder Agent converts the decompilation results into MCP format, extracting information such as function signatures, parameter semantics, calling conventions, and security clues (such as strings and magic numbers) to construct MCP messages.
[0147] 2. Test driver generates Agent: Generates test driver harness based on exported tables / interface functions.
[0148] 3. Sample generation Agent: Generate samples based on program semantics.
[0149] 4. Identify the sample file format (sample generation agent) and combine it with common file type rules, such as Pdf / png / jpeg, to load file format rules and guide sample generation / mutation (mutation strategy).
[0150] Specifically, if the MCP contains file format information (such as "PNG") for the target function, the sample generation agent will load a preset file type rule base and generate or mutate test samples according to the rules.
[0151] 5. Based on the identified format and platform (binary program analysis agent), obtain the target program's system platform information (Windows / Linux / macOS). In the test environment, deploy the agent to select different system Docker containers according to the target, compile Harness / execute Fuzzer.
[0152] Specifically, based on the target program platform information (Windows / Linux / macOS), the test environment deployment agent will select the corresponding Docker image. For example, the Windows environment image is used to compile the Windows PE format harness and perform fuzz testing on that platform.
[0153] 6. Validation of the Agent: Combining ① Testcase test samples, ② Harness fuzz test drivers, ③ the test environment obtained by deploying the Agent in the test environment, and ④ dynamic debugging (gdb / windbg / lldb) and dynamic instrumentation (Pin / DynamoRIO / Frida / tinyinst) tools, perform error analysis to verify whether the sample data reaches the target program. If the sample does not trigger the target code, regenerate the harness; if the sample triggers the target code / inputs the harness...
[0154] Specifically, dynamic debugging tools (such as gdb, windbg, lldb) and dynamic instrumentation tools (such as Pin, DynamoRIO, Frida, tinyinst) are used to verify in the test container whether the test sample can correctly trigger the target function and execute along the expected path. If the verification fails, the agent will adjust the harness generation strategy and regenerate until the verification is passed.
[0155] 7. Fuzz Agent, containerized fuzz testing (Alf++ / libfuzzer), dynamic instrumentation.
[0156] Specifically, after passing the validity verification, the Fuzz Agent will perform mutation tests in the corresponding test container using fuzzing tools (such as AFL++ or libfuzzer). Dynamic instrumentation technology will record the program's execution path, guide the mutation strategy, and improve test coverage.
[0157] 8. Feedback-driven test optimization, coverage-guided mutation is repeatedly executed (used to optimize test samples).
[0158] Specifically, the agent analyzes edge coverage statistics, performs fine-grained mutations on test samples, and prioritizes testing paths with low coverage. This strategy optimizes the efficiency of fuzz testing and ensures that as much code as possible is tested.
[0159] 9. Vulnerability analysis agent, combined with crash, outputs a report.
[0160] Specifically, when a crash or abnormal behavior is discovered during fuzzing, the vulnerability analysis agent will conduct in-depth analysis of these anomalies to verify whether they are real vulnerabilities, assess the exploitability and risk level of the vulnerabilities, and finally generate a detailed vulnerability report, which may include vulnerability description, location, scope of impact, and remediation recommendations.
[0161] In this system, multiple intelligent agents form a closed-loop, multi-stage automated fuzz testing system, which can perform efficient and comprehensive security assessments of closed-source binary programs. The process involves several stages: The binary program analysis agent first analyzes the closed-source binary program using a decompilation tool. The MCPBuilder agent then transforms these decompiled results into structured, typed, and semantically clear MCP messages. The test driver generation agent generates a test driver (Harness) that can correctly load and drive the closed-source program for fuzzing, based on the semantic information of the target function described in the MCP message. The sample generation agent uses the semantic information in the MCP message and a pre-defined file type rule base to generate structured test samples that conform to the expected format of the target program. The test environment deployment agent selects and configures the appropriate Docker container based on the platform information of the binary program provided in the MCP message, for subsequent Harness compilation, test sample execution, and fuzzing. The validity verification agent compiles the Harness in the test container and runs it for the first time, using dynamic taint analysis technology to verify whether the test sample can correctly guide the execution of the target program and whether the parameters flow through the expected path. Once the Harness passes the validity verification, the Fuzz agent executes fuzzing using the test sample in the test container, recording the execution path through dynamic instrumentation (such as AFL++ or LibFuzzer) to discover new vulnerabilities. The vulnerability analysis agent processes the fuzz... The agent captures abnormal behavior, performs Proof-of-Concept (PoC) validity verification, assesses vulnerability risks, and ultimately generates a vulnerability analysis report.
[0162] The aforementioned system architecture integrates multiple stages, including decompilation, AI Agent, dynamic verification, fuzzing, and vulnerability analysis, to achieve a fully automated fuzzing process for closed-source binary programs. The introduction of intelligent agents solves the challenges of generating test drivers and mutating test samples without source code, while the closed-loop feedback mechanism further optimizes the accuracy and efficiency of the tests. This makes the entire fuzzing process not only highly automated but also capable of deeply uncovering potential security issues in the software.
[0163] It should be noted that, Figure 3 The system shown is used to perform fuzz testing. Figure 2 The fuzzing method shown is therefore Figure 2The relevant explanations in the fuzz testing method also apply to Figure 3 The fuzzing system shown will not be described in detail here.
[0164] This application's embodiments extract program semantics through decompilation and encapsulate them into a Model Context Protocol (MCP), which is understood by the AIAgent to automatically generate test drivers and structured test cases. Combined with dynamic taint analysis to verify effectiveness, it ultimately achieves end-to-end automated fuzz testing. This technical solution achieves the following technical effects:
[0165] (1) MCP as a standardized bridge between decompiled semantics and AI: It structures, types, and semanticizes the decompiled output (such as the pseudocode of Ghidra / IDA) into a context protocol that LLM can accurately understand, solving the core problem that "AI cannot understand binary semantics". Traditional LLM directly processes undefined4 FUN_123(...) which leads to a high illusion rate, while MCP explicitly passes function signatures, parameter semantics, calling conventions, and security clues (such as strings and magic numbers), which can significantly improve the accuracy of AI-generated harnesses.
[0166] (2) Achieve a closed loop of “end-to-end fully automatic closed-source program fuzzing”: Connect the entire process of “decompilation → semantic extraction → AI generation → dynamic verification → containerized fuzz → PoC generation” without manual intervention. Traditional solutions require security researchers to manually write harnesses (which takes several hours / functions), while this solution achieves automatic coverage in minutes, which is especially suitable for firmware, drivers, commercial software and other scenarios without source code.
[0167] (3) The “multi-specialized AI Agent + MCP communication” architecture decomposes the fuzzing test task into specialized agents such as binary program analysis, test driver generation, sample generation, test environment deployment, validity verification, and fuzzing. Through MCP message collaboration, it realizes a highly cohesive, loosely coupled, and scalable intelligent agent system. Unlike the simple “single call LLM” mode, this architecture supports failure rollback, result feedback, and dynamic retry (such as harness verification failure → automatic correction and regeneration of prompt), forming an autonomous optimization closed loop.
[0168] (4) "AI-driven protocol-aware sample generation" mechanism: using semantic clues in MCP (such as the "PNG" string, the magic number 0x89504E47), the input format is automatically inferred and structured initial samples are generated in combination with the file type rule base. Compared with traditional AFL and other tools, the efficiency of structured inputs such as PNG / JSON / PDF is extremely low (>99% of samples are filtered by early verification).
[0169] (5) "Cross-platform containerized fuzzing" adaptive execution: Automatically identifies binary format (ELF / PE / Mach-O / ), dynamically selects the corresponding OS Docker image (Windows / Linux / macOS), solves the "harness compilation environment mismatch" problem, realizes a unified fuzzing pipeline for Windows PE / Linux ELF / macOS Mach-O, eliminates the need for manual configuration of the compilation environment, and greatly improves the feasibility of industrial deployment.
[0170] (6) Closed-loop feedback architecture of MCP-AI-verification-fuzz testing: Through lightweight dynamic taint analysis, taints are injected during the first execution to verify whether the objective function is called correctly and whether the parameters flow through the expected path, thus avoiding invalid tests. Based on the validity verification results or crash information of harness, the MCP extraction strategy or AI prompt is optimized in reverse, and harness is repeatedly generated until the verification is passed.
[0171] It should be noted that, for fuzz testing of closed-source programs, in addition to the aforementioned technical solutions based on Model Context Protocol (MCP) and AIAgent, the following alternative solutions can be adopted depending on the technical approach, tool selection, and application scenario:
[0172] 1) MCP format alternative: Protocol Buffers or a custom binary format can be used to replace JSON, improving transmission efficiency; incremental MCP (transmitting only the changed parts) is supported, making it suitable for large-scale program analysis. It should be noted that its readability is poor, as the binary format cannot be directly read or debugged manually; flexibility is reduced, as Protobuf requires schema modification and recompilation; incremental MCP implementation is complex, requiring the design of change detection mechanisms (such as AST diff), version number management, and merge conflict handling, significantly increasing engineering complexity.
[0173] 2) AI Agent Alternatives: Using rule engines + template filling (such as the UEFI harness template library based on EDK II); deploying small models (such as TinyLLM) on edge devices to achieve lightweight harness generation. Its advantages are high determinism, no illusions, fast execution speed, extremely low resource consumption, and strong domain adaptability, but its generalization ability is weak, and rule engines cannot handle unknown patterns (such as new protocol parsers); development and maintenance costs are high, and a completely new template needs to be written for each new platform (such as VxWorks, QNX); it lacks flexibility, and templates are difficult to express complex logic (such as "if the function calls `malloc`, then `free` needs to be added to the harness"), while LLM can make dynamic inferences; small models have limited capabilities, and models <1B have insufficient understanding of complex MCPs (such as nested structures, cross-functional dependencies), resulting in significantly lower generation quality than large models.
[0174] 3) Verification mechanism alternatives: Pre-verify the harness logic using static dataflow analysis (based on MCP); combine this with symbolic execution (such as Angr) to verify parameter reachability. The advantages are: static analysis is fast and has no execution overhead; symbolic execution provides accurate verification and can generate PoCs. However, due to its static nature, its accuracy is limited; symbolic execution suffers from poor performance and path explosion; and it cannot detect runtime behavior, such as runtime issues like "the harness calls a function but the firmware refuses to write due to SPI locking" that cannot be detected by static / symbolic methods.
[0175] Figure 4 This is a structural diagram of a fuzz testing apparatus according to an embodiment of this application, as shown below. Figure 4 As shown, the device includes:
[0176] The decompilation module 402 is used to decompile the target program to obtain the semantic information of the target program. The target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program.
[0177] Encoding module 404 is used to encode semantic information into a target message that conforms to the context protocol;
[0178] Analysis module 406 is used to analyze the target message using an intelligent agent to obtain the test driver and test samples;
[0179] Test module 408 is used to perform fuzz testing on the target program using a test driver and test samples.
[0180] It should be noted that, Figure 4 The apparatus shown is used to perform fuzz testing. Figure 2 The fuzzing method shown is therefore Figure 2The relevant explanations in the fuzz testing method also apply to Figure 4 The apparatus for fuzz testing shown will not be described in detail here.
[0181] This application also provides an electronic device, which includes a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the steps of implementing the fuzz testing method in various embodiments of this application.
[0182] This application also provides a non-volatile storage medium including a stored computer program, wherein the device containing the non-volatile storage medium executes the steps of the fuzz testing method in various embodiments of this application by running the computer program.
[0183] This application also provides a computer program product, including computer instructions that, when executed by a processor, implement the steps of the fuzz testing method in various embodiments of this application.
[0184] This application also provides a computer program that, when executed by a processor, implements the steps of the fuzz testing method in various embodiments of this application.
[0185] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0186] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0187] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0188] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0189] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0190] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0191] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for fuzz testing, characterized in that, include: The target program is decompiled to obtain its semantic information, wherein the target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program; The semantic information is encoded into a target message that conforms to the context protocol; An intelligent agent is used to analyze the target message to obtain a test driver and test samples; The target program is fuzz-tested using the test driver and the test sample.
2. The method according to claim 1, characterized in that, The target program is decompiled to obtain its semantic information, including: The target program is decompiled to obtain first semantic information and second semantic information, wherein the first semantic information includes function metadata for guiding the construction of the test driver, and the second semantic information includes a data parsing mechanism for guiding the generation of test samples; The first semantic information and the second semantic information are used as the semantic information.
3. The method according to claim 2, characterized in that, Encoding the semantic information into a target message conforming to the context protocol includes: The semantic information is encoded using a preset format to obtain the target message, wherein the target message includes a Model Context Protocol (MCP) message, which is used to explicitly convey the binary semantics of the target program.
4. The method according to claim 1, characterized in that, The intelligent agent includes a first intelligent agent and a second intelligent agent; the intelligent agents analyze the target message to obtain a test driver and test samples, including: The first intelligent agent processes the target message and the first preset prompt word to obtain the test driver program, wherein the first preset prompt word is used to guide the first intelligent agent to generate executable code of the test driver program corresponding to the target message; The second intelligent agent processes the target message and the second preset prompt word to obtain the test sample, wherein the second preset prompt word is used to guide the second intelligent agent to generate test sample metadata corresponding to the target message.
5. The method according to claim 4, characterized in that, The second intelligent agent processes the target message and the second preset prompt word to obtain the test sample, including: The target message and the second preset prompt word are used as input to the second intelligent agent to obtain test sample metadata; The file format in the test sample metadata is compared with the preset file format; If the file format matches the preset file format, the second agent is driven to select a preset rule corresponding to the preset file format from a preset structured rule base to generate the test sample; If the file format does not match the preset file format, the second agent is driven to generate a test sample corresponding to the binary semantics of the target message.
6. The method according to claim 1, characterized in that, The method further includes: Extract the platform information of the target program from the target message; Identify container images that match the platform information, wherein the container images include images corresponding to multiple systems respectively; Configure a test container corresponding to the container image, wherein the test container is used to compile the test driver.
7. The method according to claim 6, characterized in that, After configuring the test container corresponding to the container image, the method further includes: The test driver is compiled in the test container to obtain the compiled target test driver. A marker is set on the target data field of the test sample to obtain the target test sample. The marker is used to verify whether the target function of the target program is called correctly and whether the parameters passed by the test driver to the target function flow through the expected path. The target test sample is injected into the input parameters of the target test driver, the target test driver with the injected input parameters is run, and a verification result is obtained, wherein the verification result is used to indicate the validity of the test driver.
8. The method according to claim 7, characterized in that, After obtaining the verification results, the method further includes: If the verification result indicates that the verification failed, the test driver corresponding to the target program is regenerated according to the error type in the verification result. If the verification result indicates successful verification, the target program is fuzz-tested using the test driver and the test sample.
9. A fuzz testing apparatus, characterized in that, include: A decompilation module is used to decompile a target program to obtain the semantic information of the target program, wherein the target program includes a closed-source binary program, and the semantic information is used to characterize the data processing logic of the target program; The encoding module is used to encode the semantic information into a target message that conforms to the context protocol; The analysis module is used to analyze the target message using an intelligent agent to obtain the test driver and test samples; The testing module is used to perform fuzz testing on the target program using the test driver and the test samples.
10. An electronic device, characterized in that, include: A memory and a processor, the memory being used to store program instructions; the processor being connected to the memory and used to execute a method for implementing the fuzz test according to any one of claims 1 to 8.
11. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, wherein the device containing the non-volatile storage medium executes the fuzz testing method according to any one of claims 1 to 8 by running the computer program.
12. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the fuzz testing method according to any one of claims 1 to 8.