Implementation method and system for adapting an existing fuzz testing engine to a parallel architecture
By designing parallel interactive interfaces and adaptation layers, the problem of resource waste and cross-platform support of the fuzz testing engine in the parallelization process is solved, efficient fuzz testing is achieved, multilingual compatibility and fine-grained data synchronization is supported, and testing efficiency is improved.
Patent Information
- Application Number
- CN202510559478.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the process of parallelization, the existing fuzz testing engines have problems such as wasting resources, lack of effective synchronization mechanisms for collaborative work, lack of cross-platform support and limited testing efficiency improvement.
Design a parallel interactive interface, build communication channels through interface definition language and gRPC framework, create adaptation layers, decouple the fuzz testing engine and parallel logic, realize efficient data transmission and processing, support multilingual compatibility, and parallel node registration, status reporting, seed upload and logout operations.
It improves the efficiency of fuzz testing, supports cross-platform parallelization, simplifies the parallelization adaptation process of the engine, realizes fine-grained data synchronization and resource management, and improves the overall testing efficiency.
Smart Images

Figure CN120086148B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer security technology, and in particular relates to an implementation method and system for adapting an existing fuzzy testing engine to a parallelized architecture. Background Art
[0002] Fuzz testing is an automated dynamic testing technology that focuses on discovering vulnerabilities in various software under test (including applications, databases, operating systems, etc.) to improve the security and robustness of the software. By inputting randomly generated invalid or unexpected data into the program and dynamically monitoring whether the execution of the target program triggers program exceptions, fuzz testing can effectively trigger potential errors and vulnerabilities and help developers identify security risks. With the use of various program analysis techniques, such as control flow analysis, data flow analysis, symbolic execution, taint analysis, etc., a large number of fuzz testing projects have been spawned.
[0003] However, with the increasing complexity of software functions and structures, traditional single-task mode fuzz testing can no longer meet the needs of obtaining test results in real time. Therefore, researchers have gradually turned to parallel fuzz testing methods to improve the testing efficiency per unit time. Representative projects include Google's Oss-Fuzz, Microsoft's OneFuzz, CollabFuzz, EnFuzz, etc. These methods usually rely on the parallel functions of existing fuzz testing engines (such as AFL, LibFuzzer, honggfuzz) to achieve test collaboration among multiple parallel nodes through shared seed directories.
[0004] Although parallel fuzz testing has improved efficiency to a certain extent, it still faces many challenges. First, the collaborative work between different fuzz testing engines lacks an effective synchronization mechanism, which may lead to resource waste and reduced testing efficiency. Second, existing fuzz testing engines are usually optimized for specific platforms (such as specific operating systems or programming languages), lacking versatility and cross-platform support. Third, most parallelization schemes are limited to a few basic fuzz testing engines, and the selection of fuzz testing engines often relies on simple and rough heuristic rules, resulting in limited room for improvement in testing efficiency and resource utilization. Although there are endless studies on improving fuzz testing performance, how to effectively integrate multiple fuzz testing engines for parallel fuzz testing remains an urgent problem to be solved.
[0005] Faced with this challenge, it is crucial to build an efficient parallel architecture that can adapt to the existing fuzz testing engines for parallelization. Summary of the invention
[0006] In view of the above technical problems, the present invention provides an implementation method and system for adapting an existing fuzzy testing engine to a parallelized architecture.
[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0008] A method for adapting an existing fuzz testing engine to a parallel architecture, the method comprising the following steps:
[0009] S100: Design a parallel interaction interface required for parallel operation of the fuzz testing engine using an interface definition language;
[0010] S200: Design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages;
[0011] S300: Create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure decoupling between the fuzz testing engine and the parallel logic;
[0012] S400: Deeply analyze the source code of the fuzz testing engine to be adapted, and determine the insertion points for the interaction operations of the adaptation layer to ensure the smooth execution of parallel operations;
[0013] S500: Conduct multi-angle verification on the adapted fuzz testing engine to ensure data interaction with the adaptation layer without affecting the original functions, so as to implement the fuzz testing engine after parallel adaptation.
[0014] Preferably, the interaction interfaces in S200 include a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface;
[0015] The registration interface is used to register new fuzz testing nodes to the server or other fuzz testing nodes. The input parameters of the registration interface include the Fuzzer structure and the Node structure. Among them, Fuzzer is a structure that defines the information of the current fuzz testing engine, and Node is a structure that runs the environment of the current fuzz testing engine; the Fuzzer structure includes the unique identifier id of the model testing engine, the type type of the fuzz testing engine, the execution times exec, the current execution times present_exec, the coverage metric standard metrics of the current fuzz testing engine, the size metrics_value measured by the current fuzz testing engine, the registration time start_time, and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of node CPU cores cores, the CPU running percentage cpu_usage_percentage, the available memory size memory_size of the node, and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzzer ID fuzzer_id assigned within the parallel framework, and the node ID node_id assigned within the parallel framework.
[0016] The status report interface is used to periodically send the status of the fuzz testing node to the server or other fuzz testing nodes; the input parameters of the status report interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the execution times exec, the current execution times present_exec, the current time current_time, the CPU running percentage cpu_usage_percentage, and the available memory size memory_size of the node. The output parameter of the status report interface is the request status status.
[0017] The seed upload interface is used to synchronize seeds, that is, test cases that cover new paths or trigger new crashes, to the server or other fuzz testing nodes; the input parameters of the seed upload interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the Seed structure that records the seed-related information, and the trace_map that records the coverage information of the current seed. Among them, the Seed structure includes the seed ID id assigned within the parallel framework, the seed type type, the length or size length of the seed, the fuzzer ID fuzzer_id that discovered the seed, the name file_path of the seed, and the seed data data. The output parameters of the seed upload interface include the request status status.
[0018] The seed synchronization interface is used to synchronize the seeds of other fuzz testing nodes or the server to the current fuzz testing node; the input parameters of the seed synchronization interface include the fuzzer ID fuzzer_id allocated within the parallel framework, the ID number sync_seed_id of the seeds to start synchronizing, and the number of seeds to synchronize sync_seed_num. The output parameters of the seed synchronization interface include the request status status, the Seed recording the seed-related information, and the trace_map recording the coverage information of the current seeds.
[0019] The logout interface is used to log out the current fuzz testing node from the parallel task. The input parameter of the logout interface includes the fuzzer ID fuzzer_id allocated within the parallel framework, and the output parameter of the logout interface includes the request status status.
[0020] Preferably, S300 includes:
[0021] S310: Build a communication channel using the gRPC framework, implement multiplexed stream transmission using the HTTP / 2 protocol, and define standard interfaces through the Interface Definition Language Protocol Buffers to ensure that fuzz testing nodes can perform parallel communication operations efficiently and seamlessly;
[0022] S320: Use the protoc compiler to generate data structure class libraries for various programming language versions from the proto file, implement data format conversion between different programming languages, improve compatibility between multiple languages, and ensure that different fuzz testing engines can understand, process, and convert the transmitted data;
[0023] S330: Provide data interfaces for different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer.
[0024] Preferably, S400 includes:
[0025] S410: Add a parallelization option to the startup parameters of the fuzz testing engine to be adapted to support the startup of parallel operations; during the initialization or configuration phase of the engine, insert a call to the node registration interface to send out the initial information of the engine to ensure the correct registration of parallel nodes;
[0026] S420: Insert a status reporting interface into the code segment for updating and displaying status information in the main loop of the fuzz testing engine to be adapted, for real-time reporting of the test status; among them, the code segment for updating and displaying status information is a function called in a loop or periodically, used to collect the current status data and format and output it to the terminal or log.
[0027] S430: Locate and generate seeds, discover new coverage, or save the corresponding function or code segment location in the main loop of the fuzz testing engine to be adapted with parallel design, and insert a call to the seed upload interface; for an engine without parallel design, insert a call to the seed upload interface when no new paths or crash seeds are found during periodic checks or within a unit of time; where the call to the seed upload interface includes the seeds for new paths / crashes and the corresponding coverage information.
[0028] S440: Call the seed synchronization interface in the existing synchronization code segment of the fuzz testing engine to be adapted or periodically. The fuzz testing engine triggers seed synchronization at multiple stages, such as after new seeds occur, periodic synchronization, and when the test state reaches a threshold. Whether to synchronize coverage information during seed synchronization depends on the cost of synchronizing coverage information and the cost of the client re - executing to generate coverage. If the cost of synchronizing coverage information is greater than the cost of the client re - executing, the client is selected to re - execute the seed to obtain coverage information; otherwise, the coverage information is synchronized together when synchronizing the seeds.
[0029] S450: When the fuzz testing engine to be adapted ends the loop and enters the exit phase, insert a call to the logout interface to end the parallel task.
[0030] Preferably, the process of implementing data interaction with the adaptation layer in S500 is specifically as follows:
[0031] The client sends a registration request to the server or other fuzz testing nodes through the registration interface of the adaptation layer. The registration request contains the client's unique identifier and node running environment information. After receiving the registration request, the server or other fuzz testing nodes allocate a unique node ID to the client and update or add to the node status table.
[0032] The client regularly sends a status report request to the server or other fuzz testing nodes through the status report interface of the adaptation layer, including the execution status of the current node. The server or other fuzz testing nodes confirm the receipt of the status report and update the node status table.
[0033] The client submits seeds and the basic block information of their execution to the server or other fuzz testing nodes through the seed upload interface of the adaptation layer. The server or other fuzz testing nodes confirm the receipt of the seeds and update the seed queue and global basic block information.
[0034] The client requests to synchronize the latest seed information through the seed synchronization interface of the adaptation layer, including the seed ID of the previous synchronization and the current node ID. The server or other fuzz testing nodes return the currently synchronized seeds and the basic block information of their execution.
[0035] The client sends a logout request to the server or other fuzz testing nodes through the logout interface of the adaptation layer, requesting to be removed from the system; the server or other fuzz testing nodes confirm the logout request and update the node status table.
[0036] Throughout the process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and the seed queue are maintained by the server for tracking the status of each node and the seed information.
[0037] Preferably, the parallel interaction interface is applicable to the feedback-based gray box / white box fuzz testing engine, and is also applicable to the black box fuzz testing or hybrid fuzz testing without any feedback and guidance.
[0038] Preferably, the parallel interaction interface is used for data synchronization, and is also used to implement resource warning and health detection mechanisms according to the status report interface, and to implement timed scheduling updates, performance-guided scheduling updates, overhead-guided performance scheduling updates, and periodic scheduling update policies based on the seed synchronization interface.
[0039] Preferably, the parallel interaction interface synchronizes data between nodes in a synchronous, asynchronous, multi-threaded or multi-process manner to meet different performance requirements; the design of the parallel interaction interface is applicable to various architecture modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave and hybrid architecture.
[0040] Preferably, the method further includes:
[0041] Construct a central database based on the parallel interaction interface, and use servers, cloud service platforms and other methods to manage seeds to achieve more efficient seed management and resource scheduling.
[0042] Adapt the existing fuzz testing engine to an implementation system with a parallelized architecture, including a parallel interaction interface design module, a data structure improvement module, an adaptation layer creation module, an insertion point determination module and a verification module;
[0043] The parallel interaction interface design module is used to design the parallel interaction interface required for the parallel operation of the fuzz testing engine using the interface definition language.
[0044] The data structure improvement module is used to design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages.
[0045] The adaptation layer creation module is used to create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure the decoupling between the fuzz testing engine and the parallel logic.
[0046] An insertion point determination module, which is used to deeply analyze the source code of the fuzz testing engine to be adapted, determine the insertion points for the interaction operations of the adaptation layer, and ensure the smooth execution of parallel operations;
[0047] A verification module, which is used to verify the adapted fuzz testing engine from multiple perspectives, ensure data interaction with the adaptation layer without affecting the original functions, and thus implement the fuzz testing engine after parallel adaptation.
[0048] The above implementation method and system for adapting an existing fuzz testing engine into a parallel architecture integrates parallel functions into the code of the existing fuzz testing engine through the design of an adaptation layer and a standardized interface, reduces the overhead caused by file retrieval and repeated execution of seeds, significantly improves the efficiency of fuzz testing, and simplifies the parallel adaptation process of the existing fuzz testing engine. Through this method, the fuzz testing engine can be parallel across platforms and support finer-grained data synchronization and resource management, thereby improving the overall efficiency of parallel fuzz testing. Description of the Drawings
[0049] Figure 1 It is a flowchart of an implementation method for adapting an existing fuzz testing engine into a parallel architecture according to an embodiment of the present invention;
[0050] Figure 2 It is an overall architecture diagram according to an embodiment of the present invention;
[0051] Figure 3 It is a schematic diagram of a peer-to-peer network structure applicable to an embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of a master-slave network structure applicable to an embodiment of the present invention. Detailed Embodiments
[0053] In order to enable those skilled in the art of this technology to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0054] In one embodiment, as Figure 1 shown, an implementation method for adapting an existing fuzz testing engine into a parallel architecture, the method includes the following steps:
[0055] S100: Use an Interface Definition Language to design parallel interaction interfaces required for parallel operations of the fuzz testing engine; further, use an Interface Definition Language (IDL), such as Protocol Buffers (.proto file), JSON, XML, etc. to define parallel interaction interfaces;
[0056] S200: Design and improve the data structure required for the parallel interaction interface based on the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient message transmission and processing.
[0057] S300: Create an adaptation layer to encapsulate all parallel-related operations into standardized interfaces to ensure decoupling between the fuzz testing engine and the parallel logic. Further, gRPC can be used to implement this adaptation layer. gRPC will generate corresponding code for different programming languages according to the defined.proto file. For example, for the C++ language, gRPC will generate.pb.h and.pb.cc files, and for the Python language, it will generate corresponding Python files. These generated codes will enable the interaction interface to be supported in different languages, simplifying the communication between different fuzz testing engines.
[0058] S400: Deeply analyze the source code of the fuzz testing engine to be adapted to determine the insertion points for the interaction operations of the adaptation layer to ensure the smooth execution of parallel operations.
[0059] S500: Conduct multi-angle verification on the adapted fuzz testing engine to ensure data interaction with the adaptation layer without affecting the original functions, thereby implementing the fuzz testing engine after parallel adaptation.
[0060] In one embodiment, the interaction interface in S200 includes a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface.
[0061] The registration interface is used to register new fuzz testing nodes to the server or other fuzz testing nodes. The input parameters of the registration interface include the Fuzzer structure and the Node structure. Among them, Fuzzer is a structure that defines the information of the current fuzz testing engine, and Node is a structure that runs the environment of the current fuzz testing engine; the Fuzzer structure includes the unique identifier id of the model testing engine, the type type of the fuzz testing engine, the execution times exec, the current execution times present_exec, the coverage metric standard metrics of the current fuzz testing engine, the size metrics_value measured by the current fuzz testing engine, the registration time start_time, and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of node CPU cores cores, the CPU running percentage cpu_usage_percentage, the available memory size memory_size of the node, and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzzer number fuzzer_id allocated within the parallel framework, and the node number node_id allocated within the parallel framework.
[0062] The status report interface is used to periodically send the status of the fuzz testing node to the server or other fuzz testing nodes; the input parameters of the status report interface include the fuzzer number fuzzer_id allocated within the parallel framework, the execution times exec, the current execution times present_exec, the current time current_time, the CPU running percentage cpu_usage_percentage, and the available memory size memory_size of the node. The output parameter of the status report interface is the request status status.
[0063] The seed upload interface is used to synchronize seeds, that is, test cases that cover new paths or trigger new crashes, to the server or other fuzz testing nodes; the input parameters of the seed upload interface include the fuzzer number fuzzer_id allocated within the parallel framework, the Seed structure that records the seed-related information, and the trace_map that records the coverage information of the current seed. Among them, the Seed structure includes the seed number id allocated within the parallel framework, the seed type type, the length or size length of the seed, the fuzzer number fuzzer_id that discovered the seed, the name file_path of the seed, and the seed data data. The output parameters of the seed upload interface include the request status status.
[0064] The seed synchronization interface is used to synchronize seeds from other fuzz testing nodes or the server to the current fuzz testing node; the input parameters of the seed synchronization interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the ID number sync_seed_id at which to start synchronizing seeds, and the number of seeds to synchronize sync_seed_num, and the output parameters of the seed synchronization interface include the request status status, the Seed recording seed-related information, and the trace_map recording the coverage information of the current seeds;
[0065] The logout interface is used to log out the current fuzz testing node from the parallel task. The input parameter of the logout interface includes the fuzzer ID fuzzer_id assigned within the parallel framework, and the output parameter of the logout interface includes the request status status.
[0066] Specifically, the registration interface: is used to register a new fuzz testing node to the server or other fuzz testing nodes;
[0067] The input parameters of the registration interface: include the Fuzzer structure and the Node structure;
[0068] Fuzzer is a structure defining the information of the current fuzz testing engine, and Node is a structure running the environment of the current fuzz testing engine;
[0069] The general data structure is as follows:
[0070] Fuzzer{
[0071] id (integer): The unique identifier of the fuzz testing engine, used to identify and locate the registered fuzz testing node;
[0072] type (enumeration type): The type of the fuzz testing engine;
[0073] exec (integer): The number of executions;
[0074] present_exec (integer): The current number of executions;
[0075] metrics (string): The coverage metrics of the current fuzz testing engine, such as basic blocks, edges, lines, etc.;
[0076] metrics_value (integer): The size measured by the current fuzz testing engine;
[0077] start_time (timestamp): The registration time;
[0078] current_time (timestamp): The current time;
[0079] }
[0080] Node{
[0081] id (integer): Node number;
[0082] ip_address (string): Node IP address;
[0083] cores (integer): Number of CPU cores of the node;
[0084] cpu_usage_percentage (float): CPU running percentage;
[0085] memory_size (integer): Available memory of the node, used for health detection;
[0086] output (string): Output directory of the node;
[0087] }
[0088] The output parameters of the registration interface are as follows:
[0089] status (byte): Request status, for subsequent processing based on the status;
[0090] fuzzer_id (integer): Fuzzer number allocated within the parallel framework;
[0091] node_id (integer): Node number allocated within the parallel framework.
[0092] The status report interface is used to periodically send the fuzz testing node status to the server or other fuzz testing nodes;
[0093] The input parameters of the status report interface are as follows:
[0094] fuzzer_id (integer): Fuzzer number allocated within the parallel framework;
[0095] exec (integer): Number of executions;
[0096] present_exec (integer): Current number of executions;
[0097] current_time (timestamp): Current time;
[0098] cpu_usage_percentage (float): CPU running percentage;
[0099] memory_size (integer): Available memory of the node, used for health detection;
[0100] The output parameters of the status report interface are as follows:
[0101] status (bytes): The request status, and subsequent processing is based on the status.
[0102] The seed upload interface is used to synchronize "interesting" seeds (i.e., test cases that cover new paths or trigger new crashes) to the server or other fuzz testing nodes;
[0103] The input parameters of the seed upload interface are as follows:
[0104] fuzzer_id (integer): The fuzz tester number assigned within the parallel framework;
[0105] Seed (structure): Records information related to the seed;
[0106] trace_map (structure): Records the coverage information of the current seed; Note: The representation of coverage information may vary for different fuzz testing engines;
[0107] Seed{
[0108] id (integer): The seed number assigned within the parallel framework;
[0109] type (enumeration value): The seed type, such as normal seed, new seed, crash seed, timeout seed, etc.;
[0110] length (integer): The length or size of the seed;
[0111] fuzzer_id (integer): The fuzz tester number assigned within the parallel framework, the fuzz tester that discovered the seed;
[0112] file_path (string): The name of the seed in the current fuzz tester, facilitating location, statistical analysis;
[0113] data (bytes): The seed data;
[0114] }
[0115] The output parameters of the seed upload interface are as follows:
[0116] status (bytes): The request status, and subsequent processing is based on the status;
[0117] The seed synchronization interface: Used to synchronize seeds from other fuzz testing nodes or the server to the current fuzz testing node;
[0118] The input parameters of the seed synchronization interface:
[0119] fuzzer_id (integer): The fuzzer number allocated within the parallel framework;
[0120] sync_seed_id: The id number at which the synchronized seeds should start;
[0121] sync_seed_num: The number of seeds to be synchronized;
[0122] Output parameters of the seed synchronization interface:
[0123] status (byte): The request status, based on which subsequent processing is carried out;
[0124] Seed (structure): Records information related to the seeds, the synchronized seeds;
[0125] trace_map (structure): Records the coverage information of the current seeds, the coverage information corresponding to the synchronized seeds;
[0126] The cancellation interface is used to cancel the current fuzzing node from the parallel task;
[0127] Input parameters of the cancellation interface:
[0128] fuzzer_id (integer): The fuzzer number allocated within the parallel framework;
[0129] The output parameters of the cancellation interface are as follows:
[0130] status (byte): The request status, based on which subsequent processing is carried out.
[0131] In one embodiment, S300 includes:
[0132] S310: Build a communication channel using the gRPC framework, implement multiplexed stream transmission using the HTTP / 2 protocol, and define standard interfaces through the Interface Definition Language Protocol Buffers to ensure that fuzzing nodes can perform parallel communication operations efficiently and seamlessly;
[0133] S320: Use the protoc compiler to generate data structure libraries for various programming language versions from the proto file, implement data format conversion between different programming languages, improve cross-language compatibility, to ensure that different fuzzing engines can understand, process, and convert the transmitted data;
[0134] S330: Provide data interfaces for different programming languages to ensure that fuzzing engines written in different programming languages can perform parallel communication operations through the adaptation layer.
[0135] In one embodiment, S400 includes:
[0136] S410: Add a parallelization option to the startup parameters of the fuzz testing engine to be adapted to support the startup of parallel operations. Further, it also includes adding parameters related to parallelism, such as the number of parallel nodes, synchronization time interval, etc., in order to flexibly configure and optimize resource usage; during the initialization or configuration phase of the engine, insert a call to the node registration interface to send out the initial information of the engine to ensure the correct registration of parallel nodes;
[0137] S420: Insert a status reporting interface in the code segment for updating and displaying status information in the main loop of the fuzz testing engine to be adapted, for real-time reporting of the test status; among them, the code segment for updating and displaying status information is a function called in a loop or periodically, which is used to collect the current status data and format it for output to the terminal or log;
[0138] S430: Insert a call to the seed upload interface at the position of the function or code segment for positioning, seed generation, discovering new coverage, or saving in the main loop of the fuzz testing engine to be adapted with parallel design; for an engine without parallel design, insert a call to the seed upload interface when no new paths or crash seeds are found during periodic checks or within a unit time; among them, the call to the seed upload interface includes the seeds for discovering new paths / crashes and the corresponding coverage information;
[0139] S440: Call the seed synchronization interface in the existing synchronization code segment of the fuzz testing engine to be adapted or periodically; the fuzz testing engine triggers seed synchronization at multiple stages, such as after a new seed occurs, periodic synchronization, and when the test status reaches a threshold; whether to synchronize coverage information during seed synchronization depends on the overhead of synchronizing coverage information and the overhead of the client re-executing to generate coverage. If the overhead of synchronizing coverage information is greater than the overhead of the client re-executing, then choose to let the client re-execute the seed to obtain coverage information; otherwise, synchronize the coverage information together when synchronizing the seeds;
[0140] S450: Insert a call to the logout interface when the fuzz testing engine to be adapted ends the loop and enters the exit phase to end the parallel task. Further, when the fuzz testing engine to be adapted exits or crashes and enters the program before exiting, insert a call to the logout interface to ensure that the parallel task can end normally.
[0141] In one embodiment, as Figure 2 shown, the process of implementing data interaction with the adaptation layer in S500 is specifically as follows:
[0142] The client sends a registration request to the server or other fuzz testing nodes through the registration interface of the adaptation layer. The registration request contains the unique identifier of the client and the node running environment information. After receiving the registration request, the server or other fuzz testing nodes allocate a unique node ID to the client and update or add to the node status table.
[0143] The client regularly sends a status report request to the server or other fuzz testing nodes through the status report interface of the adaptation layer, including the execution status of the current node. The server or other fuzz testing nodes confirm the receipt of the status report and update the node status table.
[0144] The client submits the seed and its executed basic block information to the server or other fuzz testing nodes through the seed upload interface of the adaptation layer. The server or other fuzz testing nodes confirm the receipt of the seed and update the seed queue and the global basic block information.
[0145] The client requests to synchronize the latest seed information through the seed synchronization interface of the adaptation layer, including the seed ID of the previous synchronization and the current node ID. The server or other fuzz testing nodes return the currently synchronized seeds and their executed basic block information.
[0146] The client sends a logout request to the server or other fuzz testing nodes through the logout interface of the adaptation layer, requesting to be removed from the system. The server or other fuzz testing nodes confirm the logout request and update the node status table.
[0147] In the whole process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and the seed queue are maintained by the server, which are used to track the status of each node and the seed information.
[0148] In one embodiment, the parallel interaction interface is applicable to feedback-based grey-box / white-box fuzz testing engines (such as fuzz testing based on coverage feedback, taint analysis, symbolic execution), and is also applicable to black-box fuzz testing or hybrid fuzz testing without any feedback and guidance (such as synchronizing the seeds found by the grey-box fuzz testing engine into the black-box fuzz testing to discover potential paths).
[0149] In one embodiment, the parallel interaction interface is used for data synchronization, and is also used to implement resource warning and health detection mechanisms according to the status report interface, and to implement timed scheduling updates, performance-guided scheduling updates, overhead-guided performance scheduling updates, and periodically scheduled update strategies based on the seed synchronization interface.
[0150] In one embodiment, the parallel interaction interface synchronizes data between nodes in a synchronous, asynchronous, multi-threaded, or multi-process manner to meet different performance requirements; the design of the parallel interaction interface is applicable to multiple architecture modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave, and hybrid architecture.
[0151] In one embodiment, the method further includes:
[0152] Construct a central database based on the parallel interaction interface, and use servers, cloud service platforms, and other methods to manage seeds to achieve more efficient seed management and resource scheduling.
[0153] In one embodiment, after multi-angle verification of the adapted fuzz testing engine, including function verification of data interfaces, data integrity checks, performance tests, verification of anomaly detection capabilities, compatibility tests, and stress tests, etc., to ensure the successful completion of the adaptation process; further, the adapted fuzz testing engine can support multiple parallel modes, such as single-machine multi-process parallel, distributed parallel, and heterogeneous collaborative parallel, etc.
[0154] Compared with the prior art, the beneficial effects of the present invention are:
[0155] 1. The biggest advantage of the present invention is the improvement of fuzz testing efficiency. By adding a parallelization function to the fuzz testing engine, the present invention integrates the interaction operations between fuzz testing nodes and other nodes or servers into the adaptation layer; this design decouples the existing fuzz tester, realizes parallel fuzz testing with fine-grained data synchronization, avoids the transmission of full-scale seeds or coverage data, thereby reducing data redundancy and network overhead, and significantly improving the overall efficiency of fuzz testing;
[0156] 2. Through the design of the adaptation layer, the present invention can easily integrate fuzz testing engines in different programming languages and platforms. The use of standardized interfaces further simplifies the parallelization adaptation process of existing fuzz testing engines, improves the flexibility and scalability of the system, and can quickly adapt to changing test requirements.
[0157] In one embodiment, there is also provided an implementation system for adapting an existing fuzz testing engine to a parallel architecture, including a parallel interaction interface design module, a data structure improvement module, an adaptation layer creation module, an insertion point determination module, and a verification module;
[0158] The parallel interaction interface design module is used to design the parallel interaction interface required for the parallel operation of the fuzz testing engine using the interface definition language;
[0159] A data structure improvement module, which is used to design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages;
[0160] An adaptation layer creation module, which is used to create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure decoupling between the fuzz testing engine and the parallel logic;
[0161] An insertion point determination module, which is used to deeply analyze the source code of the fuzz testing engine to be adapted, determine the insertion points of the interaction operations of the adaptation layer, and ensure the smooth execution of parallel operations;
[0162] A verification module, which is used to verify the adapted fuzz testing engine from multiple perspectives, ensure data interaction with the adaptation layer without affecting the original functions, and thus implement the fuzz testing engine after parallel adaptation.
[0163] For the specific limitations of the implementation system for adapting the existing fuzz testing engine to a parallel architecture, reference can be made to the limitations of the implementation method for adapting the existing fuzz testing engine to a parallel architecture in the above text, which will not be elaborated here. Each module in the above implementation system for adapting the existing fuzz testing engine to a parallel architecture can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0164] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the implementation method for adapting the existing fuzz testing engine to a parallel architecture are realized.
[0165] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the implementation method for adapting the existing fuzz testing engine to a parallel architecture are realized.
[0166] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0167] The above has introduced in detail the implementation methods and systems for adapting an existing fuzz testing engine to a parallel architecture provided by the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only for helping to understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An implementation method for adapting an existing fuzz testing engine to a parallel architecture, characterized in that, The method includes the following steps: S100: Design parallel interaction interfaces required for the parallel operation of the fuzz testing engine using Interface Definition Language; S200: Design and improve the data structures required for the parallel interaction interfaces according to the interaction interfaces and the existing data structures of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; the interaction interfaces include a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface; S300: Create an adaptation layer to encapsulate all parallel-related operations into standardized interfaces to ensure decoupling between the fuzz testing engine and the parallel logic; S300 includes: S310: Build a communication channel using the gRPC framework, implement multiplexed stream transmission using the HTTP / 2 protocol, and define standard interfaces through the Interface Definition Language Protocol Buffers to ensure that fuzz testing nodes can perform parallel communication operations efficiently and seamlessly; S320: Use the protoc compiler to generate data structure class libraries in various programming language versions from proto files, implement data format conversion between different programming languages, improve cross-language compatibility, and ensure that different fuzz testing engines can understand, process, and convert the transmitted data; S330: Provide data interfaces for different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer; S400: Deeply analyze the source code of the fuzz testing engine to be adapted to determine the insertion points for the interaction operations of the adaptation layer to ensure the smooth execution of parallel operations; S500: Conduct multi-angle verification on the adapted fuzz testing engine to ensure data interaction with the adaptation layer without affecting the original functions, so as to implement the fuzz testing engine after parallel adaptation.
2. The method according to claim 1, characterized in that, The registration interface in S200 is used to register new fuzz testing nodes to the server or other fuzz testing nodes. The input parameters of the registration interface include the Fuzzer structure and the Node structure. Among them, Fuzzer is a structure that defines the information of the current fuzz testing engine, and Node is a structure that runs the environment of the current fuzz testing engine; the Fuzzer structure includes the unique identifier id of the model testing engine, the type type of the fuzz testing engine, the execution times exec, the current execution times present_exec, the coverage measurement standard metrics of the current fuzz testing engine, the size metrics_value measured by the current fuzz testing engine, the registration time start_time, and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of CPU cores cores of the node, the CPU running percentage cpu_usage_percentage, the available memory size memory_size of the node, and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzzer ID fuzzer_id allocated within the parallel framework, and the node ID node_id allocated within the parallel framework. The status report interface is used to periodically send the status of the fuzz testing node to the server or other fuzz testing nodes; the input parameters of the status report interface include the fuzzer ID fuzzer_id allocated within the parallel framework, the execution times exec, the current execution times present_exec, the current time current_time, the CPU running percentage cpu_usage_percentage, and the available memory size memory_size of the node. The output parameter of the status report interface is the request status status. The seed upload interface is used to synchronize seeds, that is, test cases that cover new paths or trigger new crashes, to the server or other fuzz testing nodes; the input parameters of the seed upload interface include the fuzzer ID fuzzer_id allocated within the parallel framework, the Seed structure that records the relevant information of the seed, and the trace_map that records the coverage information of the current seed. Among them, the Seed structure includes the seed ID id allocated within the parallel framework, the seed type type, the length or size length of the seed, the fuzzer ID fuzzer_id that discovers the seed, the name file_path of the seed, and the seed data data. The output parameter of the seed upload interface includes the request status status. The seed synchronization interface is used to synchronize the seeds of other fuzzing nodes or the server to the current fuzzing node; the input parameters of the seed synchronization interface include the fuzzer ID assigned within the parallel framework, the sync_seed_id of the seed to start synchronizing, and the number of seeds to synchronize sync_seed_num. The output parameters of the seed synchronization interface include the request status status, the Seed recording the seed-related information, and the trace_map recording the coverage information of the current seed. The logout interface is used to log out the current fuzzing node from the parallel task. The input parameter of the logout interface includes the fuzzer ID assigned within the parallel framework, and the output parameter of the logout interface includes the request status status.
3. The method according to claim 2, wherein S400 includes: S410: Add a parallelization option to the startup parameters of the fuzzing engine to be adapted to support the startup of parallel operations; during the initialization or configuration phase of the engine, insert a call to the node registration interface to send out the initial information of the engine to ensure the correct registration of parallel nodes. S420: Insert a status report interface in the code segment for updating and displaying status information in the main loop of the fuzzing engine to be adapted, for real-time reporting of the test status; among them, the code segment for updating and displaying status information is a function called in a loop or periodically, used to collect the current status data and format and output it to the terminal or log. S430: Insert a call to the seed upload interface at the position of the function or code segment for positioning, seed generation, discovery of new coverage, or saving in the main loop of the fuzzing engine to be adapted with parallel design; for an engine without parallel design, insert a call to the seed upload interface when no new paths or crash seeds are found during periodic checks or within a unit time; among them, the call to the seed upload interface includes the discovery of new path / crash seeds and the corresponding coverage information. S440: Call the seed synchronization interface in the existing synchronization code segment of the fuzzing engine to be adapted or periodically. The seed synchronization of the fuzzing engine is triggered at multiple stages after the occurrence of new seeds, periodic synchronization, and when the test status reaches a threshold; whether to synchronize the coverage information during seed synchronization depends on the overhead of synchronizing the coverage information and the overhead of the client re-executing to generate coverage. If the overhead of synchronizing the coverage information is greater than the overhead of the client re-executing, then choose to let the client re-execute the seed to obtain the coverage information; otherwise, synchronize the coverage information together when synchronizing the seeds. S450: Insert a call to the logout interface when the fuzzing engine to be adapted ends the loop and enters the exit phase to end the parallel task.
4. The method according to claim 3, characterized in that, The specific process of implementing data interaction with the adaptation layer in S500 is as follows: The client sends a registration request to the server or other fuzzing nodes through the registration interface of the adaptation layer. The registration request contains the client's unique identifier and node running environment information; after receiving the registration request, the server or other fuzzing nodes assign a unique node ID to the client and update or add the node status table. The client regularly sends status report requests to the server or other fuzz testing nodes through the status report interface of the adaptation layer, including the execution status of the current node; the server or other fuzz testing nodes confirm the receipt of the status report and update the node status table; The client submits seeds and their executed basic block information to the server or other fuzz testing nodes through the seed upload interface of the adaptation layer; the server or other fuzz testing nodes confirm the receipt of the seeds and update the seed queue and global basic block information; The client requests to synchronize the latest seed information through the seed synchronization interface of the adaptation layer, including the seed ID of the previous synchronization and the current node ID; the server or other fuzz testing nodes return the currently synchronized seeds and their executed basic block information; The client sends a logout request to the server or other fuzz testing nodes through the logout interface of the adaptation layer, requesting to be removed from the system; the server or other fuzz testing nodes confirm the logout request and update the node status table; In the whole process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and the seed queue are maintained by the server, which are used to track the status of each node and the seed information.
5. The method according to claim 4, wherein The parallel interaction interface is applicable to the feedback-based gray box / white box fuzz testing engine, and is also applicable to the black box fuzz testing or hybrid fuzz testing without any feedback and guidance.
6. The method according to claim 5, wherein The parallel interaction interface is used for data synchronization, and is also used to implement resource warning and health detection mechanisms according to the status report interface, and to implement timed scheduling updates, performance-guided scheduling updates, overhead-guided performance scheduling updates, and periodic scheduling update strategies based on the seed synchronization interface.
7. The method according to claim 6, wherein The parallel interaction interface uses synchronous, asynchronous, multi-threaded or multi-process methods to perform data synchronization between nodes to meet different performance requirements; the design of the parallel interaction interface is applicable to a variety of architecture modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave and hybrid architecture.
8. The method according to claim 7, wherein The method further includes: Constructing a central database based on the parallel interaction interface, and using servers, cloud service platforms and other methods to manage seeds to achieve more efficient seed management and resource scheduling.
9. An implementation system for adapting an existing fuzz testing engine to a parallelized architecture, characterized in that, Including a parallel interaction interface design module, a data structure improvement module, an adaptation layer creation module, an insertion point determination module and a verification module; The parallel interaction interface design module is used to design the parallel interaction interface required for the parallel operation of the fuzz testing engine using the interface definition language; The data structure improvement module is used to design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; The interaction interface includes a registration interface, a status report interface, a seed upload interface, a seed synchronization interface and a logout interface; The adaptation layer creation module is used to create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure the decoupling between the fuzz testing engine and the parallel logic; The adaptation layer creation module includes: building a communication channel using the gRPC framework, implementing multiplexed stream transmission using the HTTP / 2 protocol, and defining a standard interface through the Interface Definition Language Protocol Buffers to ensure that the fuzz testing nodes can perform parallel communication operations efficiently and seamlessly; using the protoc compiler to generate data structure libraries for various programming language versions from proto files, implementing data format conversion between different programming languages, improving compatibility between multiple languages, and ensuring that different fuzz testing engines can understand, process, and convert the transmitted data; providing data interfaces for different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer. The insertion point determination module is used to deeply analyze the source code of the fuzz testing engine to be adapted, determine the insertion points for the interaction operations of the adaptation layer, and ensure the smooth execution of parallel operations. The verification module is used to verify the adapted fuzz testing engine from multiple perspectives, ensure data interaction with the adaptation layer without affecting the original functions, and thus implement the fuzz testing engine after parallel adaptation.
Citation Information
Patent Citations
Network security test integrated device
CN113055408A
Intelligent fuzzy test method and test system for network protocol software
CN115543823A