Implementation method and system for adapting existing fuzzy test engine to parallelization architecture
By designing parallel interaction interfaces and adaptation layers, the existing fuzz testing engines have solved the problem of collaborative work and cross-platform support in parallelization, and efficient parallel adaptation of fuzz testing has been achieved, which significantly improves testing efficiency and resource utilization.
Patent Information
- Application Number
- CN202510559478.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing fuzz testing engines have a lack of synchronization mechanism for synergistic work, a lack of universality and cross-platform support, and relying on heuristic rules when selecting engines, resulting in limited room for improvement in testing efficiency and resource utilization.
By designing parallel interaction interfaces, improving data structures, creating adaptation layers and determining insertion points, and performing multi-angle verification of the fuzz testing engine, parallel adaptation of the existing fuzz testing engine is achieved.
It significantly improves the efficiency of fuzz testing, realizes cross-platform parallelism, supports finer-grained data synchronization and resource management, and simplifies the parallel adaptation process of the fuzz testing engine.
Smart Images

Figure CN120086148A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer security technology, and in particular relates to an implementation method and system for adapting an existing fuzzy testing engine to a parallelized architecture. Background Art
[0002] Fuzz testing is an automated dynamic testing technology that focuses on discovering vulnerabilities in various software under test (including applications, databases, operating systems, etc.) to improve the security and robustness of the software. By inputting randomly generated invalid or unexpected data into the program and dynamically monitoring whether the execution of the target program triggers program exceptions, fuzz testing can effectively trigger potential errors and vulnerabilities and help developers identify security risks. With the use of various program analysis techniques, such as control flow analysis, data flow analysis, symbolic execution, taint analysis, etc., a large number of fuzz testing projects have been spawned.
[0003] However, with the increasing complexity of software functions and structures, traditional single-task mode fuzz testing can no longer meet the needs of obtaining test results in real time. Therefore, researchers have gradually turned to parallel fuzz testing methods to improve the testing efficiency per unit time. Representative projects include Google's Oss-Fuzz, Microsoft's OneFuzz, CollabFuzz, EnFuzz, etc. These methods usually rely on the parallel functions of existing fuzz testing engines (such as AFL, LibFuzzer, honggfuzz) to achieve test collaboration among multiple parallel nodes through shared seed directories.
[0004] Although parallel fuzz testing has improved efficiency to a certain extent, it still faces many challenges. First, the collaborative work between different fuzz testing engines lacks an effective synchronization mechanism, which may lead to resource waste and reduced testing efficiency. Second, existing fuzz testing engines are usually optimized for specific platforms (such as specific operating systems or programming languages), lacking versatility and cross-platform support. Third, most parallelization schemes are limited to a few basic fuzz testing engines, and the selection of fuzz testing engines often relies on simple and rough heuristic rules, resulting in limited room for improvement in testing efficiency and resource utilization. Although there are endless studies on improving fuzz testing performance, how to effectively integrate multiple fuzz testing engines for parallel fuzz testing remains an urgent problem to be solved.
[0005] Faced with this challenge, it is crucial to build an efficient parallel architecture that can adapt to the existing fuzz testing engines for parallelization. Summary of the invention
[0006] In view of the above technical problems, the present invention provides an implementation method and system for adapting an existing fuzzy testing engine to a parallelized architecture.
[0007] The technical solution adopted by the present invention to solve its technical problems is as follows: A method for adapting an existing fuzz testing engine to a parallel architecture, the method comprising the following steps: S100: Design a parallel interaction interface required for parallel operation of the fuzz testing engine using an interface definition language; S200: Design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; S300: Create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure decoupling between the fuzz testing engine and the parallel logic; S400: Deeply analyze the source code of the fuzz testing engine to be adapted, and determine the insertion point of the interaction operation of the adaptation layer to ensure the smooth execution of parallel operations; S500: Verify the adapted fuzz testing engine from multiple perspectives to ensure data interaction with the adaptation layer without affecting the original functions, thereby implementing the fuzz testing engine after parallel adaptation.
[0008] Preferably, the interaction interfaces in S200 include a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface; The registration interface is used to register a new fuzz testing node to the server or other fuzz testing nodes. The input parameters of the registration interface include a Fuzzer structure and a Node structure. Among them, Fuzzer is a structure defining the information of the current fuzz testing engine, and Node is a structure running the environment of the current fuzz testing engine; the Fuzzer structure includes the unique identifier id of the model testing engine, the type type of the fuzz testing engine, the execution times exec, the current execution times present_exec, the coverage measurement standard metrics of the current fuzz testing engine, the size metrics_value measured by the current fuzz testing engine, the registration time start_time, and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of CPU cores cores of the node, the CPU running percentage cpu_usage_percentage, the available memory size memory_size of the node, and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzz tester number fuzzer_id assigned within the parallel framework, and the node number node_id assigned within the parallel framework; The status report interface is used to periodically send the status of the fuzz testing node to the server or other fuzz testing nodes; the input parameters of the status report interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the number of executions exec, the current execution count present_exec, the current time current_time, the CPU usage percentage cpu_usage_percentage, and the memory size available to the node. The output parameter of the status report interface is the request status status; The seed upload interface is used to synchronize seeds, i.e., test cases that cover new paths or trigger new crashes, to the server or other fuzz testing nodes; the input parameters of the seed upload interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the Seed structure that records seed-related information, and the trace_map that records the coverage information of the current seed. Among them, the Seed structure includes the seed ID id assigned within the parallel framework, the seed type type, the length or size length of the seed, the fuzzer ID fuzzer_id that discovered the seed, the name of the seed file_path, and the seed data data. The output parameters of the seed upload interface include the request status status; The seed synchronization interface is used to synchronize seeds from other fuzz testing nodes or the server to the current fuzz testing node; the input parameters of the seed synchronization interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the id number sync_seed_id at which the seeds to be synchronized start, and the number of seeds to be synchronized sync_seed_num. The output parameters of the seed synchronization interface include the request status status, the Seed that records seed-related information, and the trace_map that records the coverage information of the current seed; The logout interface is used to log out the current fuzz testing node from the parallel task. The input parameter of the logout interface includes the fuzzer ID fuzzer_id assigned within the parallel framework, and the output parameter of the logout interface includes the request status status.
[0009] Preferably, S300 includes: S310: Build a communication channel using the gRPC framework, implement multiplexed stream transmission using the HTTP / 2 protocol, and define standard interfaces through the Interface Definition Language Protocol Buffers to ensure that fuzz testing nodes can perform parallel communication operations efficiently and seamlessly; S320: Use the protoc compiler to generate data structure libraries for various programming language versions from the proto file, implement data format conversion between different programming languages, improve compatibility between multiple languages, and ensure that different fuzz testing engines can understand, process, and convert the transmitted data; S330: Provide data interfaces for different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer.
[0010] Preferably, S400 includes: S410: Add a parallelization option to the startup parameters of the fuzz testing engine to be adapted to support the startup of parallel operations; during the initialization or configuration phase of the engine, insert a call to the node registration interface to send out the initial information of the engine to ensure the correct registration of parallel nodes; S420: Insert a status reporting interface in the code segment for updating and displaying status information in the main loop of the fuzz testing engine to be adapted for real-time reporting of the test status; among them, the code segment for updating and displaying status information is a function called in a loop or periodically, which is used to collect current status data and format and output it to the terminal or log; S430: Insert a call to the seed upload interface at the position of the function or code segment for positioning, seed generation, discovery of new coverage, or saving in the main loop of the fuzz testing engine to be adapted with parallel design; for an engine without parallel design, insert a call to the seed upload interface when no new paths or crash seeds are found during periodic checks or within a unit time; among them, the call to the seed upload interface includes the seeds for discovering new paths / crashes and the corresponding coverage information; S440: Call the seed synchronization interface in the existing synchronization code segment of the fuzz testing engine to be adapted or periodically; the fuzz testing engine triggers seed synchronization at multiple stages such as after a new seed occurs, periodic synchronization, and when the test status reaches a threshold; whether to synchronize coverage information during seed synchronization depends on the overhead of synchronizing coverage information and the overhead of the client re-executing to generate coverage. If the overhead of synchronizing coverage information is greater than the overhead of the client re-executing, then choose to let the client re-execute the seed to obtain coverage information; otherwise, synchronize the coverage information together when synchronizing the seeds; S450: Insert a call to the logout interface when the fuzz testing engine to be adapted ends the loop and enters the exit phase to end the parallel task.
[0011] Preferably, the process of implementing data interaction with the adaptation layer in S500 is specifically as follows: The client sends a registration request to the server or other fuzz testing nodes through the registration interface of the adaptation layer. The registration request contains the unique identifier of the client and the node running environment information; after receiving the registration request, the server or other fuzz testing nodes allocate a unique node ID to the client and update or add the node status table; The client regularly sends status report requests to the server or other fuzz testing nodes through the status report interface of the adaptation layer, including the execution status of the current node; the server or other fuzz testing nodes confirm the receipt of the status report and update the node status table; The client submits seeds and their executed basic block information to the server or other fuzz testing nodes through the seed upload interface of the adaptation layer; the server or other fuzz testing nodes confirm the receipt of the seeds and update the seed queue and global basic block information; The client requests to synchronize the latest seed information through the seed synchronization interface of the adaptation layer, including the seed ID of the previous synchronization and the current node ID; the server or other fuzz testing nodes return the currently synchronized seeds and their executed basic block information; The client sends a logout request to the server or other fuzz testing nodes through the logout interface of the adaptation layer, requesting to be removed from the system; the server or other fuzz testing nodes confirm the logout request and update the node status table; In the whole process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and the seed queue are maintained by the server, which are used to track the status of each node and the seed information.
[0012] Preferably, the parallel interaction interface is applicable to the feedback-based gray box / white box fuzz testing engine, and is also applicable to the black box fuzz testing or hybrid fuzz testing without any feedback and guidance.
[0013] Preferably, the parallel interaction interface is used for data synchronization, and is also used to implement resource early warning and health detection mechanisms according to the status report interface, and to implement timed scheduling updates, performance-guided scheduling updates, cost-guided performance scheduling updates, and periodic scheduling update strategies based on the seed synchronization interface.
[0014] Preferably, the parallel interaction interface uses synchronous, asynchronous, multi-threaded or multi-process methods to perform data synchronization between nodes to meet different performance requirements; the design of the parallel interaction interface is applicable to various architecture modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave and hybrid architecture.
[0015] Preferably, the method further includes: Construct a central database based on the parallel interaction interface, and use servers, cloud service platforms and other methods to manage seeds to achieve more efficient seed management and resource scheduling.
[0016] Adapt the existing fuzz testing engine to an implementation system with a parallelized architecture, including a parallel interaction interface design module, a data structure improvement module, an adaptation layer creation module, an insertion point determination module and a verification module; A parallel interaction interface design module, which is used to design the parallel interaction interface required for the parallel operation of the fuzz testing engine using the interface definition language; A data structure improvement module, which is used to design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; An adaptation layer creation module, which is used to create an adaptation layer, encapsulate all parallel-related operations into standardized interfaces, and ensure decoupling between the fuzz testing engine and the parallel logic; An insertion point determination module, which is used to deeply analyze the source code of the fuzz testing engine to be adapted and determine the insertion point of the interaction operation of the adaptation layer to ensure the smooth execution of parallel operations; A verification module, which is used to verify the adapted fuzz testing engine from multiple perspectives to ensure data interaction with the adaptation layer without affecting the original functions, so as to implement the fuzz testing engine after parallel adaptation.
[0017] The above implementation method and system for adapting an existing fuzz testing engine to a parallel architecture integrate parallel functions into the code of the existing fuzz testing engine through the design of an adaptation layer and standardized interfaces, reduce the overhead caused by file retrieval and repeated execution of seeds, significantly improve the efficiency of fuzz testing, and simplify the parallel adaptation process of the existing fuzz testing engine. Through this method, the fuzz testing engine can be parallel across platforms and support finer-grained data synchronization and resource management, thereby improving the overall efficiency of parallel fuzz testing. Brief Description of the Drawings
[0018] Figure 1 It is a flowchart of the implementation method for adapting an existing fuzz testing engine to a parallel architecture in an embodiment of the present invention; Figure 2 It is an overall architecture diagram in an embodiment of the present invention; Figure 3 It is a schematic diagram of a peer-to-peer network structure applicable in an embodiment of the present invention; Figure 4 It is a schematic diagram of a master-slave network structure applicable in an embodiment of the present invention. Detailed Embodiment
[0019] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0020] In one embodiment, as Figure 1 shown, the implementation method for adapting an existing fuzz testing engine to a parallel architecture, the method includes the following steps: S100: Design the parallel interaction interfaces required for the parallel operation of the fuzz testing engine using an Interface Definition Language; further, use an Interface Definition Language (IDL), such as Protocol Buffers (.proto file), JSON, XML, etc. to define the parallel interaction interfaces; S200: Design and improve the data structures required for the parallel interaction interfaces based on the interaction interfaces and the existing data structures of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; S300: Create an adaptation layer to encapsulate all parallel-related operations into standardized interfaces to ensure decoupling between the fuzz testing engine and the parallel logic; further, gRPC can be used to implement this adaptation layer. gRPC will generate corresponding code for different programming languages according to the defined.proto file. For example, for the C++ language, gRPC will generate.pb.h and.pb.cc files, while for the Python language, it will generate corresponding Python files. These generated codes will enable the interaction interfaces to be supported in different languages, simplifying the communication between different fuzz testing engines; S400: Deeply analyze the source code of the fuzz testing engine to be adapted to determine the insertion points for the interaction operations of the adaptation layer to ensure the smooth execution of parallel operations; S500: Conduct multi-angle verification on the adapted fuzz testing engine to ensure data interaction with the adaptation layer without affecting the original functions, thereby implementing the fuzz testing engine after parallel adaptation.
[0021] In one embodiment, the interaction interfaces in S200 include a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface; The registration interface is used to register new fuzz testing nodes to the server or other fuzz testing nodes. The input parameters of the registration interface include the Fuzzer structure and the Node structure. Among them, Fuzzer is a structure that defines the information of the current fuzz testing engine, and Node is a structure that runs the environment of the current fuzz testing engine; the Fuzzer structure includes the unique identifier id of the model testing engine, the type type of the fuzz testing engine, the execution times exec, the current execution times present_exec, the coverage measurement standard metrics of the current fuzz testing engine, the size metrics_value measured by the current fuzz testing engine, the registration time start_time, and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of CPU cores cores of the node, the CPU running percentage cpu_usage_percentage, the available memory size memory_size of the node, and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzzer ID fuzzer_id assigned within the parallel framework, and the node ID node_id assigned within the parallel framework. The status report interface is used to periodically send the status of the fuzz testing node to the server or other fuzz testing nodes; the input parameters of the status report interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the execution times exec, the current execution times present_exec, the current time current_time, the CPU running percentage cpu_usage_percentage, and the available memory size memory_size of the node. The output parameter of the status report interface is the request status status. The seed upload interface is used to synchronize seeds, that is, test cases that cover new paths or trigger new crashes, to the server or other fuzz testing nodes; the input parameters of the seed upload interface include the fuzzer ID fuzzer_id assigned within the parallel framework, the Seed structure that records the seed-related information, and the trace_map that records the coverage information of the current seed. Among them, the Seed structure includes the seed ID id assigned within the parallel framework, the seed type type, the length or size length of the seed, the fuzzer ID fuzzer_id that discovered the seed, the name file_path of the seed, and the seed data data. The output parameter of the seed upload interface includes the request status status. The seed synchronization interface is used to synchronize seeds from other fuzz testing nodes or the server to the current fuzz testing node; the input parameters of the seed synchronization interface include the fuzzer ID fuzzer_id allocated within the parallel framework, the ID number sync_seed_id at which to start synchronizing seeds, and the number of seeds to be synchronized sync_seed_num, and the output parameters of the seed synchronization interface include the request status status, the Seed recording seed-related information, and the trace_map recording the coverage information of the current seeds; The logout interface is used to log out the current fuzz testing node from the parallel task. The input parameter of the logout interface includes the fuzzer ID fuzzer_id allocated within the parallel framework, and the output parameter of the logout interface includes the request status status.
[0022] Specifically, the registration interface: is used to register a new fuzz testing node to the server or other fuzz testing nodes; The input parameters of the registration interface: include the Fuzzer structure and the Node structure; Fuzzer is a structure defining the information of the current fuzz testing engine, and Node is a structure running the environment of the current fuzz testing engine; The general data structure is as follows: Fuzzer{ id (integer): The unique identifier of the fuzz testing engine, used to identify and locate the registered fuzz testing node; type (enumeration type): The type of the fuzz testing engine; exec (integer): The number of executions; present_exec (integer): The current number of executions; metrics (string): The measurement criteria covered by the current fuzz testing engine, such as basic blocks, edges, lines, etc.; metrics_value (integer): The size measured by the current fuzz testing engine; start_time (timestamp): The registration time; current_time (timestamp): The current time; } Node{ id (integer): The node number; ip_address (string): The node IP address; cores (integer): The number of CPU cores of the node; cpu_usage_percentage (floating point): The CPU running percentage; memory_size (integer): The amount of memory available to the node, used for health checks; output (string): The output directory of the node; }
[0023] The output parameters of the registration interface are as follows: status (byte): The request status, and subsequent processing is based on the status; fuzzer_id (integer): The fuzzer number allocated within the parallel framework; node_id (integer): The node number allocated within the parallel framework.
[0024] The status report interface is used to periodically send the fuzzing node status to the server or other fuzzing nodes; The input parameters of the status report interface are as follows: fuzzer_id (integer): The fuzzer number allocated within the parallel framework; exec (integer): The number of executions; present_exec (integer): The current number of executions; current_time (timestamp): The current time; cpu_usage_percentage (float): The CPU running percentage; memory_size (integer): The amount of memory available to the node, used for health checks; The output parameters of the status report interface are as follows: status (byte): The request status, and subsequent processing is based on the status.
[0025] The seed upload interface is used to synchronize "interesting" seeds (i.e., test cases that cover new paths or trigger new crashes) to the server or other fuzzing nodes; The input parameters of the seed upload interface are as follows: fuzzer_id (integer): The fuzzer number allocated within the parallel framework; Seed (structure): Records information related to the seed; trace_map (structure): Records the coverage information of the current seed; Note: The coverage information representation may be different for different fuzzing engines; Seed{ id (integer): The seed number allocated within the parallel framework; type (enumeration value): The seed type, such as normal seed, new seed, crash seed, timeout seed, etc.; length (integer): The length or size of the seed; fuzzer_id (integer): The fuzzer number assigned within the parallel framework, the fuzzer that discovered the seed; file_path (string): The name of the seed in the current fuzzer, facilitating location and statistical analysis; data (bytes): The seed data; }
[0026] The output parameters of the seed upload interface are as follows: status (bytes): The request status, for subsequent processing based on the status; Seed synchronization interface: Used to synchronize seeds from other fuzzing nodes or the server to the current fuzzing node; Input parameters of the seed synchronization interface: fuzzer_id (integer): The fuzzer number assigned within the parallel framework; sync_seed_id: The starting id number of the seeds to be synchronized; sync_seed_num: The number of seeds to be synchronized; Output parameters of the seed synchronization interface: status (bytes): The request status, for subsequent processing based on the status; Seed (structure): Records information related to the seed, the synchronized seed; trace_map (structure): Records the coverage information of the current seed, the coverage information corresponding to the synchronized seed; The logout interface is used to log out the current fuzzing node from the parallel task; Input parameters of the logout interface: fuzzer_id (integer): The fuzzer number assigned within the parallel framework; The output parameters of the logout interface are as follows: status (bytes): The request status, for subsequent processing based on the status.
[0027] In one embodiment, S300 includes: S310: Build a communication channel using the gRPC framework, implement multiplexed stream transmission using the HTTP / 2 protocol, and define standard interfaces through the Interface Definition Language Protocol Buffers to ensure that fuzzing nodes can perform parallel communication operations efficiently and seamlessly; S320: Use the protoc compiler to generate data structure libraries in various programming language versions from the proto file, implement data format conversion between different programming languages, improve compatibility between multiple languages, and ensure that different fuzz testing engines can understand, process, and convert the transmitted data; S330: Provide data interfaces for different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer.
[0028] In one embodiment, S400 includes: S410: Add a parallelization option to the startup parameters of the fuzz testing engine to be adapted to support the startup of parallel operations. Further, it also includes adding parameters related to parallelism, such as the number of parallel nodes, synchronization time interval, etc., to flexibly configure and optimize resource usage; during the initialization or configuration phase of the engine, insert a call to the node registration interface to send the initial information of the engine to ensure correct registration of parallel nodes; S420: Insert a status reporting interface in the code segment for updating and displaying status information in the main loop of the fuzz testing engine to be adapted, for real-time reporting of the test status; among them, the code segment for updating and displaying status information is a function called in a loop or periodically, used to collect current status data and format it for output to the terminal or log; S430: Insert a call to the seed upload interface at the position of the function or code segment for location and seed generation, discovering new coverage, or saving in the main loop of the fuzz testing engine to be adapted with parallel design; for an engine without a parallel design, insert a call to the seed upload interface when periodically checking or when no new paths or crash seeds are found within a unit time; among them, the call to the seed upload interface includes discovering new paths / crash seeds and the corresponding coverage information; S440: Call the seed synchronization interface in the existing synchronization code segment of the fuzz testing engine to be adapted or periodically. The fuzz testing engine triggers seed synchronization at multiple stages, such as after a new seed occurs, periodic synchronization, and when the test status reaches a threshold; whether to synchronize coverage information during seed synchronization depends on the overhead of synchronizing coverage information and the overhead of the client re-executing to generate coverage. If the overhead of synchronizing coverage information is greater than the overhead of the client re-executing, then choose to let the client re-execute the seed to obtain coverage information; otherwise, synchronize the coverage information together when synchronizing the seeds; S450: Insert a call to the logout interface when the fuzz testing engine to be adapted exits the loop and enters the exit phase to end the parallel task. Further, when the fuzz testing engine to be adapted exits or crashes and enters the program exit, insert a call to the logout interface to ensure that the parallel task can end normally.
[0029] In one embodiment, such asFigure 2 As shown in Figure 2 , the process of implementing data interaction with the adaptation layer in S500 is specifically as follows: The client sends a registration request to the server or other fuzz testing nodes through the registration interface of the adaptation layer. The registration request contains the unique identifier of the client and the node running environment information. After receiving the registration request, the server or other fuzz testing nodes allocate a unique node ID to the client and update or add the node status table. The client regularly sends a status report request to the server or other fuzz testing nodes through the status report interface of the adaptation layer, including the execution status of the current node. The server or other fuzz testing nodes confirm the receipt of the status report and update the node status table. The client submits the seed and its executed basic block information to the server or other fuzz testing nodes through the seed upload interface of the adaptation layer. The server or other fuzz testing nodes confirm the receipt of the seed and update the seed queue and the global basic block information. The client requests to synchronize the latest seed information through the seed synchronization interface of the adaptation layer, including the seed ID of the previous synchronization and the current node ID. The server or other fuzz testing nodes return the currently synchronized seeds and their executed basic block information. The client sends a deregistration request to the server or other fuzz testing nodes through the deregistration interface of the adaptation layer, requesting to be removed from the system. The server or other fuzz testing nodes confirm the deregistration request and update the node status table. In the whole process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and the seed queue are maintained by the server, which are used to track the status of each node and the seed information.
[0030] In one embodiment, the parallel interaction interface is applicable to feedback-based grey-box / white-box fuzz testing engines (such as fuzz testing based on coverage feedback, taint analysis, and symbolic execution), and is also applicable to black-box fuzz testing or hybrid fuzz testing without any feedback and guidance (such as synchronizing the seeds found by the grey-box fuzz testing engine into the black-box fuzz testing to discover potential paths).
[0031] In one embodiment, the parallel interaction interface is used for data synchronization, and is also used to implement resource warning and health detection mechanisms according to the status report interface, and to implement timed scheduling updates, performance-guided scheduling updates, overhead-guided performance scheduling updates, and periodically scheduled update policies based on the seed synchronization interface.
[0032] In one embodiment, the parallel interaction interface synchronizes data between nodes in a synchronous, asynchronous, multi-threaded, or multi-process manner to meet different performance requirements; the design of the parallel interaction interface is applicable to multiple architecture modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave, and hybrid architecture.
[0033] In one embodiment, the method further includes: Construct a central database based on the parallel interaction interface, and use servers, cloud service platforms, and other methods to manage seeds to achieve more efficient seed management and resource scheduling.
[0034] In one embodiment, after multi-angle verification of the adapted fuzz testing engine, including functional verification of data interfaces, data integrity checks, performance tests, verification of anomaly detection capabilities, compatibility tests, stress tests, etc., to ensure the successful completion of the adaptation process; further, the adapted fuzz testing engine can support multiple parallel modes, such as single-machine multi-process parallel, distributed parallel, and heterogeneous collaborative parallel, etc.
[0035] Compared with the prior art, the beneficial effects of the present invention are: 1. The biggest advantage of the present invention is the improvement of fuzz testing efficiency. By adding a parallelization function to the fuzz testing engine, the present invention integrates the interaction operations between fuzz testing nodes and other nodes or the server side into the adaptation layer; this design decouples the existing fuzz tester, realizes parallel fuzz testing with fine-grained data synchronization, avoids the transmission of full-scale seeds or coverage data, thereby reducing data redundancy and network overhead, and significantly improving the overall efficiency of fuzz testing; 2. Through the design of the adaptation layer, the present invention can easily integrate fuzz testing engines in different programming languages and platforms. The use of standardized interfaces further simplifies the parallelization adaptation process of the existing fuzz testing engine, improves the flexibility and expandability of the system, and can quickly adapt to changing test requirements.
[0036] In one embodiment, there is also provided an implementation system for adapting an existing fuzz testing engine to a parallelized architecture, including a parallel interaction interface design module, a data structure improvement module, an adaptation layer creation module, an insertion point determination module, and a verification module; The parallel interaction interface design module is used to design the parallel interaction interface required for the parallel operation of the fuzz testing engine using interface definition language; The data structure improvement module is used to design and improve the data structure required for the parallel interaction interface according to the interaction interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; An adaptation layer creation module, which is used to create an adaptation layer, encapsulate all parallel-related operations into a standardized interface, and ensure the decoupling between the fuzz testing engine and the parallel logic; An insertion point determination module, which is used to deeply analyze the source code of the fuzz testing engine to be adapted, determine the insertion points of the adaptation layer interaction operations, and ensure the smooth execution of parallel operations; A verification module, which is used to verify the adapted fuzz testing engine from multiple perspectives, ensure data interaction with the adaptation layer without affecting the original functions, and thus implement the fuzz testing engine after parallel adaptation.
[0037] For the specific limitations of the implementation system for adapting an existing fuzz testing engine to a parallel architecture, reference can be made to the limitations of the implementation method for adapting an existing fuzz testing engine to a parallel architecture in the above text, which will not be elaborated here. Each module in the above implementation system for adapting an existing fuzz testing engine to a parallel architecture can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0038] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the implementation method for adapting an existing fuzz testing engine to a parallel architecture are realized.
[0039] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the implementation method for adapting an existing fuzz testing engine to a parallel architecture are realized.
[0040] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0041] The above has introduced in detail the implementation method and system for adapting an existing fuzz testing engine to a parallel architecture provided by the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for adapting an existing fuzz testing engine to a parallel architecture, characterized in that: The method comprises the following steps: S100: Designing a parallel interaction interface required for parallel operation of the fuzz testing engine using an interface definition language; S200: Design and improve the data structure required for the parallel interactive interface according to the interactive interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; S300: Create an adaptation layer to encapsulate all parallel-related operations into standardized interfaces to ensure the decoupling between the fuzz testing engine and the parallel logic; S400: deeply analyzing the source code of the fuzz testing engine to be adapted, and determining the insertion point of the adaptation layer interaction operation to ensure the smooth execution of the parallel operation; S500: Verify the adapted fuzz test engine from multiple angles to ensure data interaction with the adaptation layer without affecting the original functions, thereby realizing the parallel adapted fuzz test engine.
2. The method according to claim 1, characterized in that The interactive interfaces in S200 include a registration interface, a status report interface, a seed upload interface, a seed synchronization interface, and a logout interface; The registration interface is used to register a new fuzz test node to the server or other fuzz test nodes. The input parameters of the registration interface include Fuzzer structure and Node structure, where Fuzzer is a structure that defines the current fuzz test engine information, and Node is a structure that runs the current fuzz test engine environment; the Fuzzer structure includes the unique identifier id of the model test engine, the type of the fuzz test engine type, the number of executions exec, the current number of executions present_exec, the coverage metrics of the current fuzz test engine metrics, the size of the current fuzz test engine metrics_value, the registration time start_time and the current time current_time; the Node structure includes the node number id, the node IP address ip_address, the number of node CPU cores cores, the CPU running percentage cpu_usage_percentage, the number of available memory memory_size of the node and the output directory output of the node; the output parameters of the registration interface include the request status status, the fuzz tester number fuzzer_id assigned internally by the parallel framework, and the node number node_id assigned internally by the parallel framework; The status report interface is used to periodically send the status of the fuzz test node to the server or other fuzz test nodes; the input parameters of the status report interface include the fuzz tester number fuzzer_id assigned by the parallel framework, the number of executions exec, the current number of executions present_exec, the current time current_time, the CPU running percentage cpu_usage_percentage and the available memory size of the node memory_size, and the output parameter of the status report interface is the request status status; The seed upload interface is used to synchronize the seed, i.e., the test case that covers the new path or triggers the new crash, to the server or other fuzz test nodes; the input parameters of the seed upload interface include the fuzz tester number fuzzer_id assigned by the parallel framework, the Seed structure that records the seed-related information, and the trace_map that records the coverage information of the current seed, wherein the Seed structure includes the seed number id assigned by the parallel framework, the seed type type, the seed length or size length, the fuzz tester number fuzzer_id that found the seed, the seed name file_path, and the seed data data; the output parameters of the seed upload interface include the request status status; The seed synchronization interface is used to synchronize the seeds of other fuzz test nodes or servers to the current fuzz test node; the input parameters of the seed synchronization interface include the fuzz tester number fuzzer_id assigned inside the parallel framework, the id number sync_seed_id of the seed to be synchronized, and the number of seeds to be synchronized sync_seed_num; the output parameters of the seed synchronization interface include the request status status, Seed that records seed-related information, and trace_map that records the coverage information of the current seed; The deregistration interface is used to deregister the current fuzz test node from the parallel task. The input parameters of the deregistration interface include the fuzz tester number fuzzer_id assigned inside the parallel framework, and the output parameters of the deregistration interface include the request status status.
3. The method according to claim 2, characterized in that S300 includes: S310: Use the gRPC framework to build a communication channel, use the HTTP / 2 protocol to implement multiplexed stream transmission, and define a standard interface through the interface definition language Protocol Buffers to ensure that fuzz test nodes can perform parallel communication operations efficiently and seamlessly; S320: Use the protoc compiler to generate data structure libraries of various programming language versions from the proto file, realize data format conversion of different programming languages, improve the compatibility between multiple languages, and ensure that different fuzz testing engines can understand, process and convert the transmitted data; S330: Provide data interfaces of different programming languages to ensure that fuzz testing engines written in different programming languages can perform parallel communication operations through the adaptation layer.
4. The method according to claim 3, characterized in that S400 includes: S410: adding a parallelization option to the startup parameters of the fuzz test engine to be adapted to support the startup of parallel operations; inserting a call to a node registration interface during the initialization or configuration phase of the engine to send out the initial information of the engine to ensure that the parallel nodes are correctly registered; S420: inserting a status reporting interface into the code segment for updating and displaying status information in the main loop of the fuzz test engine to be adapted, for real-time reporting of the test status; wherein the code segment for updating and displaying status information is a function called in a loop or at a fixed time, for collecting current status data and formatting and outputting it to a terminal or a log; S430: inserting a call to the seed upload interface during positioning and seed generation in the main loop of the fuzz test engine to be adapted with parallel design, finding new coverage or saving the corresponding function or code segment location; for the engine without parallel design, inserting a call to the seed upload interface when no new path or crash seed is found during periodic inspection or per unit time; wherein the call to the seed upload interface includes the seed for finding new path / crash and the corresponding coverage information; S440: In the existing synchronization code segment of the fuzz test engine to be adapted or the seed synchronization interface is periodically called, the fuzz test engine triggers the seed synchronization after the new seed is generated, periodic synchronization, and when the test status reaches the threshold; whether to synchronize the coverage information during seed synchronization depends on the overhead of synchronizing the coverage information and the overhead of the client re-executing the generated coverage rate. If the overhead of synchronizing the coverage information is greater than the overhead of the client re-executing, the client is selected to re-execute the seed to obtain the coverage information; otherwise, the coverage information is synchronized together with the seed synchronization; S450: When the fuzzy test engine to be adapted ends the loop and enters the exit phase, insert a call to the logout interface to end the parallel task.
5. The method according to claim 4, characterized in that The specific process of implementing data interaction with the adaptation layer in S500 is as follows: The client sends a registration request to the server or other fuzz test nodes through the registration interface of the adaptation layer. The registration request contains the client's unique identifier and the node operating environment information. After receiving the registration request, the server or other fuzz test node assigns a unique node ID to the client and updates or adds the node status table. The client periodically sends a status report request to the server or other fuzz test nodes through the status report interface of the adaptation layer, including the execution status of the current node; the server or other fuzz test nodes confirm the receipt of the status report and update the node status table; The client submits the seed and the basic block information of its execution to the server or other fuzz test nodes through the seed upload interface of the adaptation layer; the server or other fuzz test nodes confirm the receipt of the seed and update the seed queue and global basic block information; The client requests the latest seed information to be synchronized through the seed synchronization interface of the adaptation layer, including the seed ID of the last synchronization and the current node ID; the server or other fuzz testing nodes return the currently synchronized seed and the basic block information of its execution; The client sends a logout request to the server or other fuzz test nodes through the logout interface of the adaptation layer, requesting to be removed from the system; the server or other fuzz test nodes confirm the logout request and update the node status table; In the whole process, the adaptation layer is responsible for handling the communication and data conversion between the client and the server or other fuzz testing nodes. The node status table and seed queue are maintained by the server to track the status and seed information of each node.
6. The method according to claim 5, characterized in that The parallel interactive interface is suitable for feedback-based gray-box / white-box fuzz testing engines, as well as black-box fuzz testing or hybrid fuzz testing without any feedback and guidance.
7. The method according to claim 6, characterized in that The parallel interaction interface is used for data synchronization, and is also used to implement resource warning and health detection mechanisms based on the status reporting interface, as well as to implement scheduled scheduling updates, performance-guided scheduling updates, overhead-guided performance scheduling updates, and periodic scheduling update strategies based on the seed synchronization interface.
8. The method according to claim 7, characterized in that The parallel interaction interface uses synchronous, asynchronous, multi-threaded or multi-process methods to synchronize data between nodes to meet different performance requirements; the design of the parallel interaction interface is suitable for a variety of architectural modes, including peer-to-peer architecture, master-slave architecture, cascaded master-slave architecture and hybrid architecture.
9. The method according to claim 8, characterized in that The method further comprises: A central database is built based on a parallel interactive interface, and seeds are managed using servers and cloud service platforms in a variety of ways to achieve more efficient seed management and resource scheduling.
10. A system for adapting an existing fuzz testing engine to a parallelized architecture, characterized in that: It includes parallel interactive interface design module, data structure improvement module, adaptation layer creation module, insertion point determination module and verification module; A parallel interaction interface design module, used to design a parallel interaction interface required for parallel operation of a fuzz test engine using an interface definition language; The data structure improvement module is used to design and improve the data structure required by the parallel interactive interface according to the interactive interface and the existing data structure of the fuzz testing engine to be adapted, so as to ensure efficient transmission and processing of messages; The adaptation layer creation module is used to create the adaptation layer and encapsulate all parallel-related operations into standardized interfaces to ensure the decoupling between the fuzz testing engine and the parallel logic. The insertion point determination module is used to deeply analyze the source code of the fuzz test engine to be adapted and determine the insertion point of the adaptation layer interaction operation to ensure the smooth execution of parallel operations; The verification module is used to verify the adapted fuzz test engine from multiple angles to ensure data interaction with the adaptation layer without affecting the original functions, thereby realizing the parallel adapted fuzz test engine.
Citation Information
Patent Citations
Automated testing device and method of integrated heterogeneous testing tool
CN102937932A
Network security test integrated device
CN113055408A
Fuzzy testing method for embedded device program
CN115495753A
Intelligent fuzzy test method and test system for network protocol software
CN115543823A
Windows platform parallel fuzzy testing method and system based on selective instrumentation
CN118503097A