Parallel fuzzy test node optimization method and system for distributed environment

By asynchronously generating test cases and recording branch jump addresses and execution counts, using the master node to poll and merge the path array, matching duplicate paths and stopping the test process, periodically evaluating the value of seed samples, and merging high-quality seeds to low-value nodes, the problems of duplicate coverage and seed queue imbalance in distributed environments are solved, improving test efficiency and resource utilization.

CN120910863APending Publication Date: 2025-11-07NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511008623.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In a distributed environment, multi-node parallel fuzz testing suffers from issues of duplicate coverage and resource waste, and the uneven quality of seed queues among nodes leads to low testing efficiency.

Method used

Test cases are generated asynchronously and branch jump addresses and execution counts are recorded. The master node polls and merges the path array, matches duplicate paths and stops the test process. The value of seed samples is periodically evaluated, high-quality seeds are merged into low-value nodes, and the fuzzy task queue is optimized.

Benefits of technology

It reduces redundant testing, improves resource utilization and testing efficiency, achieves efficient collaboration and independence between nodes, and optimizes the seed queue distribution strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910863A_ABST
    Figure CN120910863A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of software security testing, and particularly relates to a distributed environment-oriented parallel fuzzy test node optimization method and system. The method comprises the following steps: asynchronously executing a test case generated by a fuzzy test on a first test node by using a seed sample of a fuzzy task queue, and recording a branch jump address and execution times executed in the test process to obtain a local jump path array corresponding to the first test node; utilizing the master control node to perform polling collection on the local jump path arrays of the plurality of test nodes, and executing bitwise or operation merging to obtain a global path array; utilizing the global path array and a local jump path array of the second test node to match the jump address and the execution frequency to obtain a local jump path array corresponding to the second test node; and periodically performing value evaluation and seed merging on the seed samples of the fuzzy task queue of the test node to obtain scores of the seed samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of software security testing technology, and in particular relates to a method and system for optimizing parallel fuzzy testing nodes in a distributed environment. Background Technology

[0002] In traditional single-node fuzzing environments, the incremental increase in path coverage primarily comes from the good matching of the number and depth of input mutations with lexical semantics. However, as complex software grows in scale, the processing power of a single node becomes a bottleneck, making the evolution of fuzzing towards multi-node distributed execution an inevitable choice.

[0003] However, implementing fuzz testing in a distributed environment also presents new key technical challenges. During the process of multiple nodes simultaneously generating test cases and simulating their execution in parallel, a large number of test cases are prone to "duplication coverage," meaning multiple nodes spend time repeatedly exploring paths already executed by other nodes, resulting in resource waste and decreased efficiency. The first key issue is how to synchronize path information while maintaining node independence and scalability to avoid redundant testing.

[0004] On the other hand, during fuzzing, each node maintains a local task queue containing "seed samples" for mutation. The quality of these local queues directly affects the effectiveness of subsequent test case generation by the node. Since the value of seeds varies, an unreasonable seed distribution can lead some nodes into inefficient testing or even invalid path exploration. Therefore, ensuring the basic independence of each node while optimizing the distribution of seed queues becomes the second crucial issue that urgently needs to be addressed.

[0005] In view of this, there is an urgent need for a parallel fuzz test node optimization method and system for distributed environments to solve the above-mentioned technical problems. Summary of the Invention

[0006] Therefore, it is necessary to provide a method and system for optimizing parallel fuzzy test nodes in a distributed environment to address the aforementioned technical problems.

[0007] Firstly, this application provides a method for optimizing parallel fuzzy test nodes in a distributed environment, wherein the multiple test nodes include at least a first test node and a second test node, and the method includes:

[0008] Using the seed samples of the fuzzy task queue, test cases generated by fuzz testing are executed asynchronously on the first test node, and the branch jump addresses and execution counts are recorded during the test to obtain the local jump path array corresponding to the first test node;

[0009] Collecting the local jump path arrays of the multiple test nodes by polling and performing a bitwise OR operation to obtain a global path array;

[0010] Matching the jump addresses and execution times of the global path array and the local jump path array of the second test node, and if the matching result is repeated, stopping the mutation and test process of the current seed sample, selecting a next seed sample from the local fuzzy task queue to test the second test node, and obtaining a local jump path array corresponding to the second test node;

[0011] Periodically evaluating the value of the seed sample in the fuzzy task queue of the test node to obtain a score of the seed sample;

[0012] In at least one cycle, the two test nodes with the highest and lowest scores are screened, and part of the seed samples with high scores in the test node with the highest score are selected and merged into the test node with the lowest score, and the fuzzy task queue of the test node with the lowest score is updated to obtain optimization of the test node.

[0013] In some ways, the step of using the seed sample in the fuzzy task queue to perform a fuzzy test on the first test node to generate a test case and record the branch jump address and execution times reached during the test process to obtain a local jump path array corresponding to the first test node includes:

[0014] Storing the seed sample in the fuzzy task queue in the first test node to obtain a test case generated by the fuzzy test;

[0015] Asynchronously executing the test case in the first test node and recording the branch jump address and execution times to obtain a local jump path array corresponding to the first test node.

[0016] In some ways, the step of using the master node to collect the local jump path arrays of the multiple test nodes by polling and performing a bitwise OR operation to obtain a global path array includes:

[0017] Using the master node to send a data request to the multiple test nodes according to a preset order to collect the local jump path arrays of the multiple test nodes;

[0018] Performing a bitwise OR operation on the global path array of the master node and the local jump path arrays of the multiple test nodes to obtain the global path array of the master node.

[0019] In some embodiments, the matching of the jump address and the execution times is performed by using the global path array and the local jump path array of the second test node, and if the matching result is repeated, the mutation and test process of the current seed sample of the second test node is stopped, the next seed sample is selected from the local fuzzy task queue, and the second test node is tested to obtain the local jump path array corresponding to the second test node.

[0020] The matching of the jump address is performed by using the global path array and the local jump path array of the second test node, and a jump address matching result is obtained.

[0021] If the jump address matching result is repeated, the matching of the execution times is performed, and an execution times matching result is obtained.

[0022] If the execution times matching result is repeated, the mutation and test process of the current seed sample of the second test node is stopped, the next seed sample is selected from the local fuzzy task queue, and the second test node is tested to obtain the local jump path array corresponding to the second test node.

[0023] In some embodiments, the global path array of the master node is combined with the local jump path arrays of a plurality of test nodes by performing a bitwise OR operation to obtain the global path array of the master node.

[0024] The first test node is polled and collected by using the global path array of the master node, the local jump path array of the first test node is combined into the global path array, and a new global path array is obtained.

[0025] The new global path array is distributed to the second test node, and the second test node updates its global path array copy.

[0026] In some embodiments, the value of the seed sample in the fuzzy task queue of the test node is periodically evaluated to obtain a score of the seed sample.

[0027] A seed sample value score index is constructed and a corresponding weight is assigned, wherein the seed sample value score index includes jump coverage, genetic kinship, jump danger, and sample data size.

[0028] The value of the seed sample in the fuzzy task queue of the test node is periodically evaluated according to the seed sample value score index to obtain a score of the seed sample.

[0029] The seed sample is sorted according to the score of the seed sample.

[0030] In some embodiments, the step of selecting, in the test node with the highest score, seed samples with high scores in a preset proportion and merging them into the test node with the lowest score, and updating the fuzzy task queue of the test node with the lowest score to obtain an optimized test node, comprises:

[0031] In at least one cycle, the scores of all seed samples of each test node are obtained to obtain the scores of the fuzzy task queue;

[0032] According to the scores of the fuzzy task queue, the scores of each test node are determined;

[0033] According to the scores of each test node, the two test nodes with the highest and lowest scores are obtained;

[0034] In the test node with the highest score, seed samples with high scores in a preset proportion are selected and merged into the test node with the lowest score to obtain an optimized test node.

[0035] In some embodiments, the step of using the seed samples of the fuzzy task queue to asynchronously execute the test cases generated by the fuzzy test on the first test node, and recording the branch jump addresses and execution times reached during the test to obtain the local jump path array corresponding to the first test node, comprises:

[0036] A unique ID and kinship depth information are constructed for each seed sample;

[0037] When the seed sample is mutated to form a mutant seed sample, the mutant seed sample inherits the parent seed ID and the kinship depth is increased.

[0038] In a second aspect, the application provides a parallel fuzzy test node optimization system for a distributed environment, which is applied to the parallel fuzzy test node optimization method for a distributed environment. The multiple test nodes at least include a first test node and a second test node. The system comprises:

[0039] An initial unit is configured to use the seed samples of the fuzzy task queue to asynchronously execute the test cases generated by the fuzzy test on the first test node, and record the branch jump addresses and execution times reached during the test to obtain the local jump path array corresponding to the first test node;

[0040] A polling unit is configured to use the master node to poll and collect the local jump path arrays of the multiple test nodes, and perform a bitwise OR operation to merge them to obtain a global path array;

[0041] A matching unit is configured to match the jump address and the execution times by using the global path array and the local jump path array of the second test node. If the matching result is repeated, the mutation and test process of the current seed sample are stopped, the next seed sample is selected from the local fuzzy task queue, and the second test node is tested to obtain the local jump path array corresponding to the second test node.

[0042] An evaluation unit is configured to periodically evaluate the value of the seed sample in the fuzzy task queue of the test node to obtain a score of the seed sample.

[0043] A result unit is configured to select, in at least one period, two test nodes with the highest and lowest scores, select part of the seed samples with high scores in the test node with the highest score, and merge the seed samples into the test node with the lowest score, update the fuzzy task queue of the test node with the lowest score, and obtain the optimization of the test node.

[0044] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the parallel fuzzy test node optimization method for a distributed environment when executing the computer program.

[0045] Beneficial effects: the parallel fuzzy test node optimization method for a distributed environment provided by the application comprises the following steps: a seed sample of a fuzzy task queue is utilized; test cases generated by performing fuzzy test on the first test node are recorded in an asynchronous mode; branch jump addresses and execution times reached during the test are recorded to obtain a local jump path array corresponding to the first test node; the local jump path arrays of the multiple test nodes are polled and collected by a master node, and a bitwise OR operation is performed to combine the local jump path arrays to obtain a global path array; the global path array and the local jump path array of the second test node are matched in terms of jump addresses and execution times; if the matching result is repeated, the mutation and test process of the current seed sample are stopped, the next seed sample is selected from the local fuzzy task queue to test the second test node, and a local jump path array corresponding to the second test node is obtained; the value of the seed sample of the fuzzy task queue of the test node is evaluated periodically to obtain the score of the seed sample; in at least one period, the two test nodes with the highest and lowest scores are screened, and part of the seed samples with high scores in the test node with the highest score are selected and combined into the test node with the lowest score; the fuzzy task queue of the test node with the lowest score is updated to obtain the optimization of the test node. Through the above method, the architecture burden of the traditional master node actively pushing the coverage path is bypassed, and the strategy of the distributed node self-maintaining the "jump array" is adopted, all path data is abstracted into a simple one-dimensional bit array, the master node collects in a one-way polling mode, and the information is synchronized through a simple bitwise OR combination logic. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0047] Figure 1 A flowchart of a parallel fuzzy test node optimization method for a distributed environment in an embodiment.

[0048] Figure 2 A logic diagram of a master node of a parallel fuzzy test node optimization method for a distributed environment in an embodiment. DETAILED DESCRIPTION

[0049] For the purpose of promoting the understanding of the present application, a more complete description of the application will be provided below with reference to the accompanying drawings. The embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0051] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.

[0052] Some terms related to the present application are explained below to facilitate understanding of the present application:

[0053] Fuzz testing is an automated vulnerability detection method, which randomly generates inputs in unknown or non-standard input space, tests the system's response to these unexpected inputs, and discovers potential security vulnerabilities or abnormal behaviors.

[0054] Distributed fuzz testing system refers to the allocation and parallel execution of fuzz testing tasks on multiple independent nodes (such as physical machines / containers / threads), through task division, test case sharing or result merging mechanisms, to realize the test system architecture of maximizing the efficiency of vulnerability discovery.

[0055] Jump coverage is an indicator to measure the breadth of code path execution of test cases. Jump refers to the logical branch conversion in code execution. Each unique jump path covered means a higher level of understanding of code behavior.

[0056] Path synchronization mechanism refers to the exchange of existing execution path information between different nodes in parallel fuzz testing to avoid repeated coverage of tested paths. Path synchronization is crucial to improving overall efficiency.

[0057] Queue merging optimization refers to an optimization strategy in distributed fuzz testing. Each node has a local test task queue. By evaluating the quality of local seeds, the seed sharing between nodes is dynamically adjusted to improve overall test effectiveness.

[0058] global-array, refers to an array accessible in the global scope of a program, usually used as shared data for multiple functions or modules.

[0059] total-array, usually refers to an array that stores the cumulative value of a certain calculation result (such as sum, average, etc.).

[0060] Fuzzing, as an important way of complex software security testing and vulnerability mining, has been widely used in industry, open source projects and research fields in recent years with the growth of software complexity and the increasing demand for automatic testing. LibFuzzer, AFL, Honggfuzz developed by Google, etc. all belong to such application representatives. However, traditional fuzzing is generally based on single-process and single-node model, which cannot meet the needs of large-scale system and time-sensitive security evaluation.

[0061] To improve testing efficiency, some systems attempt to implement parallel extension of fuzzing. For example:

[0062] 1. AFL's fork-server parallel mechanism

[0063] This mechanism allows testing through multi-process mechanism on the same host. Multiple worker processes obtain seeds from the main process and execute test tasks through lightweight fork mechanism to achieve parallelism. However, this mechanism is limited by local resources, does not support cluster / cross-machine deployment, and has low path synchronization granularity.

[0064] 2. Google ClusterFuzz

[0065] As an internal fuzzing cluster system of Google, ClusterFuzz can run on large-scale cloud resources, but its synchronization granularity is coarse, path sharing between nodes is inefficient, and it relies on manual plugin configuration for path coverage induction.

[0066] 3. FuzzManager and Distributed AFL

[0067] These tools achieve simple distributed control by sharing path coverage files or exchanging seed samples between nodes, but still have problems such as arbitrary queue strategy, coarse sharing granularity, and uneven task allocation.

[0068] These works provide a rudimentary architecture for parallel fuzzing, but still have three problems:

[0069] High node repetition rate leads to efficiency decline;

[0070] Path synchronization requires high-overhead daemon support;

[0071] The unreasonable distribution of seeds causes uneven node performance.

[0072] Therefore, a solution framework is needed, which has higher information aggregation efficiency, lower runtime resource consumption, and cooperation between nodes without excessive dependence on central coordination.

[0073] As described above, the existing parallel fuzzing scheme has the following defects in synchronization control mechanism and queue management strategy:

[0074] First, the path synchronization efficiency in most current distributed architectures is extremely low. The central control-based path comparison method can avoid duplication at an early stage, but as the number of nodes increases, the master node needs to synchronize thousands of coverage path files, which severely restricts the test throughput. If global path hash comparison is used, the delay caused by additional access and calculation has a great impact, and ultimately it is difficult to coordinate efficiency and accuracy.

[0075] Second, the management mechanism of the local node task queue generally lacks intelligent optimization. The node queue is often fixed and managed by the operating system or scheduling tool, and cannot be evaluated according to the seed test value, which can easily cause some nodes to be trapped in invalid path attempts, and even be ignored due to the lack of new value of the executed path, causing global redundancy to grow.

[0076] Third, since different test nodes do not share test history or knowledge accumulation, each node is prone to "repeated trial and error", and it is imagined that multiple fuzzers simultaneously blind test the same path, which wastes mutation opportunities.

[0077] The purpose of the present application is to solve the above key problems, and the goal is to:

[0078] Construct a polling path synchronization mechanism based on state space division, which greatly reduces the computational redundancy caused by repeated coverage;

[0079] Implement a local queue automatic evaluation and high-quality seed fusion algorithm, so that each node in the distributed system can maintain independent testing while also constantly absorbing test cases with higher test value;

[0080] Integrate the asynchronous message passing mechanism and the lightweight context information synchronization to ensure that the system can still run stably and efficiently in a large-scale node environment.

[0081] As shown in Figure 1 and Figure 2 , in a first aspect, the present application provides a parallel fuzzing node optimization method for a distributed environment, the method comprising:

[0082] S100, using the seed sample of the fuzzy task queue, asynchronously executing the test case generated by the fuzzy test on the first test node, and recording the branch jump address and execution times during the test process to obtain the local jump path array corresponding to the first test node.

[0083] Specifically, obtaining the local jump path array corresponding to the first test node can include the following steps:

[0084] S101, in the first test node, pre-storing the seed sample of the fuzzy task queue to obtain the test case generated by the fuzzy test.

[0085] Specifically, the first test node is only illustrative, and the seed sample of the fuzzy task queue is pre-stored in each test node. The seed sample can be an initial sample, that is, the seed samples of the test nodes are the same.

[0086] A seed sample can be extracted from the fuzzy task queue according to priority (such as queue order or weight score) to complete reading of the seed sample.

[0087] The seed sample is a historically generated test input data (such as a binary file and a protocol data packet), and the storage format can be a byte sequence.

[0088] The test node is used as a fuzzer to perform a lightweight mutation operation on the seed sample to generate a new test case:

[0089] The mutation strategy for the seed sample includes but is not limited to: bit flipping: randomly flipping bits in the seed (such as 0x01 to 0x81); byte insertion / deletion: inserting random bytes or deleting fragments in the seed; arithmetic addition / subtraction: +1 / -1 operation on the integer bytes in the seed; block replacement: replacing the seed fragment with a predefined special value (such as the boundary value 0xFFFF). Through the mutation of the seed sample, a new test case (mutated input data) is formed for subsequent execution.

[0090] S102, in the first test node, asynchronously executing the test case, and recording the branch jump address and execution times to obtain the local jump path array corresponding to the first test node.

[0091] Specifically, the execution of the plurality of test nodes can be asynchronous, and the generated test case is input into the target program (such as a parser, a kernel module, etc.) for execution. The asynchronous processing has the advantages that the node continues to process other tasks (such as queue management) and does not block waiting for the end of execution, and does not interfere with other nodes.

[0092] Need to be explained, branch jump address: record the memory address of each conditional jump / function call in the target program (such as 0x4010A2); The number of executions: record the number of times each branch jump address is triggered in this execution (such as jump address 0x4010A2 is triggered 3 times). Further, a lightweight hash table can be used to record, the format can be (jump address, execution times).

[0093] Also need to be explained, after completing the local test of the first test node, update the local jump path array, wherein in the local jump path array, total_array[] is a pre-allocated fixed-length array, the index (subscript) corresponds to the hash value of the branch jump address (index = hash (jump address) % N); Identify whether the jump address has been covered (such as 1 = covered, 0 = not covered); Execution time interval: store the number of times in the preset interval.

[0094] In this embodiment, all test nodes maintain a local jump path record array (total-array[]), which records the jump marks executed during the test. In order to facilitate subsequent running, the master node polls all test nodes in a fixed cycle order, extracts their path arrays and merges them into a global path array by performing a bitwise OR operation. The operation cost is low, and only a few CPU clock cycles are consumed for each merge.

[0095] S200, using the master node, polling and collecting the local jump path arrays of multiple test nodes, and performing a bitwise OR operation to merge and obtain a global path array.

[0096] Specifically, obtaining the global path array can include the following steps:

[0097] S201, using the master node, sending data requests to multiple test nodes according to a preset order to poll and collect local jump path arrays of multiple test nodes.

[0098] Specifically, the master node maintains a test node polling sequence (such as [test node 1, test node 2, test node 3]) and accesses each test node in a fixed order.

[0099] The preset order can be based on ascending order of test node ID or dynamic load balancing strategy (such as preferentially polling active nodes).

[0100] Next, send a lightweight data request instruction to the current target test node (through UDP / RPC communication), requiring it to return its total_array[]. The current target test node returns its local jump path array (total_array[]) asynchronously, and the data format can be a binary block. In addition, the current test node does not interrupt the local test task when responding. The request instruction can only contain control commands, without data load.

[0101] Each polling only merges the node_array[] of the current target test node, and gradually updates the global path array.

[0102] S202, perform a bitwise OR operation to merge the global path array of the master node and the local jump path array of each test node, to obtain the global path array of the master node.

[0103] Specifically, the bitwise OR operation merges the coverage state bit, traverses each element (index i corresponds to the jump address hash value) of the array, and performs a logical OR (OR) operation to ensure that any node covers the jump, and the global is marked as covered.

[0104] Next, perform the number of times merging, and the system can choose any of the following ways to process:

[0105] Cumulative mode: traverse each element in the local jump path array of the test node, and perform cumulative addition of the number of times, for example, calculate the total number of times of the jump in all nodes, and the global original number of times is 8+test node new number of times 5, and the merged result is 13. In this way, only basic arithmetic operations are required, and the calculation overhead is low.

[0106] Interval mode: update according to the preset interval (such as [1-10], [10-100]), if the number of times reported breaks the upper limit of the current interval, it will be upgraded to a higher interval (such as reporting 105 times in the original interval [10-100] will be updated to [100-1000]). In this way, the qualitative turning point of the safety sensitive jump is captured through interval breakthrough detection, and the problem of false coverage is solved.

[0107] Finally, write the merged result to the global_array[] of the master node, and update the global path array of the master node.

[0108] Each test node can be regarded as a fuzzer, which respectively performs mutation and testing, and the master node continuously polls each test node. The master node maintains a global array global-array[], which has the same element and element structure as total-array[], and is used to record global jump information.

[0109] S300, performing matching of the jump address and the execution times by using the global path array and the local jump path array of the second test node, if the matching result is repeated, stopping the mutation and test process of the current seed sample, and selecting a next seed sample in the local fuzzy task queue to test the second test node, to obtain the local jump path array corresponding to the second test node.

[0110] Wherein, once the merging is completed, the master node dispatches the updated global-array[] to the next test node, and the test node judges whether the current path has been completed by other nodes by comparing the current total-array[] with the latest global-array[], if repeated, the current mutation is immediately ended, and new path exploration is started.

[0111] The polling process is asynchronously rotated on a multi-task node system, and each test node does not need to wait for the global result, but only needs to synchronize the path state once at the appropriate rotation point. The whole mechanism saves the network bandwidth of frequent state submission, and the merging operation is simple, without recursive path tree or complete CFG comparison.

[0112] Specifically, obtaining the local jump path array corresponding to the second test node can include the following steps:

[0113] S301, performing matching of the jump address by using the global path array and the local jump path array of the second test node, to obtain a jump address matching result.

[0114] Specifically, according to the global path array issued by the master node and the local jump path array of the second test node, the jump address is matched, all jump address indexes of the second test node are traversed, for each index i, the coverage state bit of the local jump path array is checked, if local_array[i].cover_bit==1, the jump address is covered in the current test case execution, then the coverage state bit of the global path array is checked, if global_array[i].cover_bit==0, it indicates that the jump address is not covered in the global, it is judged that the jump address matching result is not repeated, and the matching process is ended.

[0115] Further, if all local covered jump addresses (i.e. all indexes i satisfying local_array[i].cover_bit==1) have been covered in the global path array (i.e. all corresponding global_array[i].cover_bit==1), it is judged that the jump address matching result is repeated, and the further judgment of S302 step is performed.

[0116] S302, if the jump address matching result is repeated, perform execution frequency matching to obtain an execution frequency matching result.

[0117] Specifically, for each index i, if local_array[i].cover_bit == 1, obtain the local execution frequency, and obtain the global execution frequency, perform execution frequency matching for each covered jump address, if the execution frequency matching result is not repeated, end the matching process, otherwise, if the execution frequencies of all covered jump addresses are the same, determine that the execution frequency matching result is repeated.

[0118] S303, if the execution frequency matching result is repeated, stop the mutation and test process of the current seed sample of the second test node, select the next seed sample in the local fuzzy task queue, and test the second test node to obtain a local jump path array corresponding to the second test node.

[0119] Specifically, after two times of jump address and execution frequency matching are determined to be repeated, the mutation and test process of the current seed sample of the second test node is stopped, that is, the mutation and test of the current seed sample have been performed on other test nodes, and there is no need to repeat the test.

[0120] It should be noted that when each new test case is generated, total-array[] is compared and updated, the master node continuously polls each test node, takes out total-array[], and combines to generate global-array[], the combination method is to perform bitwise OR operation on the corresponding elements of the array, and then update total-array[] of the next node. When comparing, first check whether all jumps have occurred, if so, then compare the number of jumps, the granularity of the number comparison is the same as when generating the test case, if they are consistent, it is considered that the current test task of the node is a repeated task, and the current test of the node is terminated and the test case is replaced.

[0121] This polling mechanism has three advantages:

[0122] First, it can achieve a head-on effect for each test node, each test node generates a new test case based on the previous work of all test nodes, reducing the occurrence of repeated tests;

[0123] Second, the polling mechanism is initiated by the master node, and each test node tests independently, reducing the computational burden, and compared with each test node actively reporting data, it can more effectively reduce the phenomenon of data conflict, especially when the system is large;

[0124] Thirdly, the master node only performs a simple bit-wise OR operation when polling, and each test node only needs to refresh a small amount of data, without stopping to wait. When the number of test nodes is large, the test time can be significantly reduced, and the polling period depends on the computing power of the master node and the number of test nodes.

[0125] In one embodiment, the step S202 of performing a bit-wise OR operation to combine the global path array of the master node and the local jump path array of each test node to obtain the global path array of the master node comprises:

[0126] S2021, polling and collecting the first test node using the global path array of the master node, and combining the local jump path array of the first test node into the global path array to obtain a new global path array.

[0127] Specifically, the master node sends a data request instruction to the first test node, for example, through UDP / RPC. The first test node returns its local jump path array asynchronously (without interrupting the local test), and uses the local jump path array to update the global path array to obtain a new global path array, which will be applied to the next test node.

[0128] S2022, distributing the new global path array to the second test node to obtain an updated global path array copy of the second test node.

[0129] Specifically, the second test node after the first test node is selected, like the first test node, performs polling and collection, and at the same time, updates the original global path array copy using the new global path array, so as to ensure that the updated global path array copy in the second test node contains the jump path in the first test node.

[0130] Specifically, the operations include the following: first, the master node initiates a polling process on the first test node according to a preset polling sequence. The polling process sends a data request instruction to the first test node; the first test node returns a local jump path array (total array) maintained by the first test node to the master node while keeping the local fuzzy test task asynchronous execution. The local jump path array records the new execution path information discovered by the first test node since the last polling, including the new covered jump address state bit (identified by 1 bit whether covered) and the corresponding execution times. After receiving the local jump path array, the master node performs a merging operation on the local jump path array and a global path array (global array) maintained by the master node: for the coverage state bit of the jump address, perform a bitwise logical OR (OR) operation to ensure that any node covered path is marked as covered globally; the system selects any of the following operations according to the configuration strategy: cumulative mode: accumulate the new number of times reported by the first test node into the global total; interval mode: update the global interval state according to the preset interval (such as [1-10], [10-100]). The merging operation generates a new global path array (new global array), which reflects the global path coverage state after the first test node is included in the latest contribution.

[0131] Next, the master node sends the newly generated new_global_array to the next node in the polling sequence, i.e., the second test node. After receiving the new_global_array (which has integrated the new path added by the first test node), the second test node performs a key update operation: overwrites its locally stored copy of the global path array (global_array_copy) with the newly received new_global_array, resets the local total_array: retains the cover status bit of the jump address (indicating path existence), and only clears the execution count, ensuring that only incremental execution data is recorded subsequently. This updated global_array_copy will be used for path deduplication decisions when generating new test cases by the node subsequently: in cumulative mode, if the current path has been completely covered in global_array_copy and the execution count reaches 90%-110% of the global value, or in interval mode, if the current path has been completely covered and the jump execution count is in the same preset interval (such as [1-10]), the current test task is terminated immediately. This global path array copy is used for path matching decisions when generating test cases by the second test node subsequently. It is worth noting that this operation only updates the global path array copy and does not overwrite or modify the second test node's private local jump path array (total_array). At the same time, to avoid repeated data statistics in subsequent polling, the second test node resets the execution count counter in its total_array: retains the cover status bit of the jump address (indicating path existence), but clears the execution count of each jump address, so that its total_array only focuses on recording the incremental execution data generated by the second test node itself in the next polling period.

[0132] This state separation and reset mechanism ensures that the newly added path information of each test node is only counted once when merged by the master node, while ensuring that the node uses the latest global path array copy for efficient path deduplication decisions. The master node continues this polling process, sequentially accessing each test node in the sequence, collecting incremental data, merging and updating the global path array, and sending updates to the next node, ultimately achieving low-overhead, high-precision synchronization of path information in a distributed environment, effectively eliminating redundant testing. Since the jump address space is fixed after initialization, the global path array (including the status bit and the execution count) always maintains a constant size; and after each polling by the node, only the jump cover status bit is retained and the execution count is cleared, ensuring that the storage overhead does not increase with the number of polling times, completely avoiding the risk of space expansion.

[0133] Need to explain, the processing of the number of executions adopts a preset interval classification mechanism instead of simple numerical accumulation. In the path synchronization process of distributed fuzz testing, when the master node merges the execution frequency data reported by the test nodes, the global state is classified and updated according to the preset jump trigger segmentation interval (such as divided into four intervals of [1-10], [10-100], [100-1000] and [1000+]).

[0134] The core operation of this mechanism is that the master node compares the execution frequency reported by the nodes with the current global recorded frequency interval in the polling process. If the reported value first breaks through the upper limit of the existing interval of a jump address (for example, the global original record is [10-100] interval, and the node reported value is 105 times), the global state of the jump address is upgraded to a higher interval (that is, updated to [100-1000]); if the reported value is within the current interval range, the original state is maintained unchanged.

[0135] This design of improving granularity through interval improves the granularity of the design, solves the problem of pseudo coverage that the path has been covered but the key behavior has not been fully triggered in the traditional scheme, and avoids the storage and synchronization overhead of accurate counting.

[0136] S400, periodically evaluate the value of the seed sample of the fuzzing task queue of the test node, and obtain the score of the seed sample.

[0137] Specifically, obtaining the score of the seed sample can include the following steps:

[0138] S401, constructing a seed sample value score index and assigning a corresponding weight.

[0139] The seed sample value score index includes jump coverage, genetic kinship, jump danger and sample data size.

[0140] Specifically, each test node periodically calculates the average test value of the "seed sample" in its detection queue, which can be scored by the following four-dimensional evaluation index:

[0141] Jump coverage (40% weight): the proportion of jumps reached by the seed;

[0142] Genetic kinship (30% weight): genetic similarity of parent / brother source;

[0143] Jump danger (20% weight): whether to cross dangerous blocks or special instructions;

[0144] Sample data size (10% weight): the more reasonable the size, the higher the score.

[0145] The calculation method of the above four indexes can be adjusted as needed, and the present application does not limit it.

[0146] S402, periodically, according to the seed sample value scoring indicator, the seed sample of the fuzzy task queue of the test node, value evaluation is carried out, and the score of the seed sample is obtained.

[0147] Specifically, the periodic trigger value evaluation, for example, automatic execution every N minutes (such as N = 30), can be started during the idle period of the node (to avoid interfering with the test task).

[0148] Traverse each seed sample in the fuzzy task queue, obtain its historical test record, including jump coverage, genetic kinship, jump danger and sample data size, use weighted summation formula to calculate, after the structure is obtained, normalization processing can be carried out according to the need, and the score of each seed sample in the fuzzy task queue is obtained.

[0149] S403, according to the score of the seed sample, the seed sample is sorted.

[0150] Specifically, the seed sample can be sorted according to the score. Next, the seed sample with a preset percentage after sorting can be selected as a high-quality seed, for example, 5%, which is used for cross-node sharing.

[0151] S500, in at least one period, the two test nodes with the highest and lowest scores are screened, and part of the seed samples with high scores in the test node with the highest score are selected and merged into the test node with the lowest score, the fuzzy task queue of the test node with the lowest score is updated, and the optimization of the test node is obtained.

[0152] Specifically, the optimization of the test node can include the following steps:

[0153] S501, in at least one period, the score of all seed samples of each test node is obtained, and the score of the fuzzy task queue is obtained.

[0154] Specifically, the cycle time can be set as needed, and in a period, the score of all seed samples of each test node is obtained, and then the scores of these seed samples are added up, for example, added or weighted calculation, to obtain the score of the fuzzy task queue.

[0155] S502, according to the score of the fuzzy task queue, the score of each test node is determined.

[0156] Specifically, after obtaining the score of the fuzzy task queue, the score is marked as the score of the test node.

[0157] S503, according to the score of each test node, the two test nodes with the highest and lowest scores are obtained.

[0158] Specifically, after marking the scores of each test node, two test nodes with the highest and lowest scores are selected as the nodes for subsequent actions according to the scores of the test nodes.

[0159] S504, in the test node with the highest score, a preset proportion of seed samples with high scores are selected and merged into the test node with the lowest score to obtain the optimization of the test node.

[0160] Specifically, in the test node with the highest score, a preset proportion of seed samples with high scores are selected, wherein the preset proportion can be 5%, 10%, etc. The selected seed samples are merged into the test node with the lowest score to form a queue merging optimization, and finally the optimization of the test node is obtained.

[0161] The preset proportion is dynamically adjusted with the period:

[0162] Obtain the lowest score of the test node;

[0163] If the lowest score is lower than the preset score value, the merging proportion of the seed sample is increased.

[0164] For example, the preset score value is 0.3, when the score of the test node with the lowest score is less than 0.3, 10% is merged, and if it is greater than or equal to 0.3, 5% is merged. Prevent the low-score test node from losing the exploration ability due to the injection of too many seeds in a short period.

[0165] It should be noted that merging into the test node with the lowest score means appending the seed sample with high score to the tail of the fuzzy task queue of the test node (not covering the original seed).

[0166] In one embodiment, in the step S100, the seed samples of the fuzzy task queue are used to perform the fuzzy test on the first test node to generate test cases, and the branch jump address and execution times reached during the test process are recorded to obtain the local jump path array corresponding to the first test node, including:

[0167] A unique ID and kinship depth information are constructed for each seed sample;

[0168] When the seed sample is mutated to form a mutated seed sample, the mutated seed sample inherits the parent seed ID and the kinship depth is increased.

[0169] Wherein, the kinship depth is 0 at the beginning, and after one mutation, the depth is increased by 1.

[0170] Specifically, the timing of the master node dispatching seeds to the test nodes, when the highest scoring seed samples are merged into the lowest scoring test nodes, specifically, the master node copies seed samples from the fuzzy task queue of the highest scoring test nodes, and adds the seed samples to the tail of the fuzzy task queue of the lowest scoring test nodes, so that the highest scoring test nodes only provide seeds (without deleting the original seeds), and the lowest scoring test nodes additionally receive seeds (without covering the original queue).

[0171] In the process of executing this process, when the master node detects that the proportion of globally high derivative depth seed samples (such as depth ≥ 5) exceeds a set threshold, such as 40%, the master node triggers a fuse mechanism, injects brand new initial seed samples (such as depth = 0), and cuts off the homologous derivative chain; the master node preferentially dispatches low depth seeds (such as 0-1) to high average depth test nodes (such as > 3), and does not deliver high derivative seeds (such as ≥ 4) to low depth test nodes (such as < 2), forcing the test nodes to break out of homologous exploration inertia; within the test nodes, such as the engine, a weight reduction is applied to the mutation seed samples with a depth ≥ 4, that is, the score is reduced according to the weight reduction on the original seed sample score, so that the computing power can be focused on the innovative exploration of low depth seeds.

[0172] Further, the master node can collect depth data of seeds of each node at each polling, and dynamically update a global depth distribution table (Hash table structure) for fuse trigger judgment.

[0173] In order to be compatible with the fuzzy device configuration module, a seed kinship tracking ID pool system is designed, each test sample contains its source ID and kinship depth information when dispatched, and depth backtracking is used to avoid "seed polling trap", that is, multiple nodes are trapped in path concentrated coverage caused by high-speed derivative homologous seeds due to common fusion.

[0174] In order to distinguish and demonstrate the advantages of the application in engineering implementation and actual effect, the current possible alternative schemes and their limitations are analyzed in detail.

[0175] Alternative scheme one: multiple nodes run independently without state, and the results are finally merged

[0176] In this method, each node does not exchange path information with others, and records crash, coverage rate and other data during testing, and does not unify the path database during testing, and only merges and compares in the final analysis stage for deduplication. This scheme seems simple, but in essence it is only a test result set, not an execution synchronization, and the result is that multiple nodes explore the same path for a long time, resulting in decreasing value.

[0177] Comparison and analysis of the application and alternative scheme one

[0178] Lack of path sharing mechanism, the same coverage path may be more than 10 times the test case repeated processing, CPU / GPU computing resources waste significantly. In contrast, the application can adjust the direction according to the path state in any round, dynamically filter repeated test tasks, and the actual test speed is improved by more than 60%.

[0179] Alternative two: synchronization system based on global path tree

[0180] The method uploads the path nodes generated by each node in the form of a complete path tree to the master end, and globally constructs a logical state diagram. When synchronizing, other nodes compare and judge uniqueness through the path tree. Although the theoretical accuracy is high, the depth of the path diagram is too long, which will cause the comparison cost to rise explosively, especially when the path hash conflict occurs, it is easy to misjudge; secondly, the path tree needs to be frequently maintained, which increases the burden of the master node.

[0181] Comparison and analysis of the application and alternative two

[0182] The application uses bit array compression to represent the jump space, and the maximum state is only a few kilobytes. In the environment of hundreds of nodes, the actual test is stable, and the synchronization delay is not more than 5ms per second.

[0183] Alternative three: use machine learning strategy to predict seed importance and automatically merge queue

[0184] Some researches try to model the path information through reinforcement learning, supervised learning and other methods, and predict the importance of seeds in different states, so as to perform automatic selective synchronization. This kind of method is strong in theoretical model explanation, but is highly sensitive to actual deployment. The training set deviation is easy to cause sampling failure, especially in the context of across-binary fuzzing, the generalization ability is limited.

[0185] Comparison and analysis of the application and alternative three

[0186] The application provides a static score and dynamic weight fusion mechanism, which can stabilize the queue structure without large data training, and can be used compatibly under any semantic model.

[0187] The algorithm of the polling mechanism is as follows:

[0188] Variable name

[0189] total-array[]

[0190] global-array[]

[0191] N

[0192] Algorithm:

[0193] Known: total-array[], global-array[], N

[0194] 1 : load global-array[] from master node

[0195] 2: if the node pool is not empty then

[0196] 3: while 1 do

[0197] 4: load total-array[] from next target task node

[0198] 5: for i <- 0 to N do

[0199] 6: global-array[i] <- global-array[i] | total-array[i]

[0200] 7: total-array[i] <- global-array[i]

[0201] 8: end for

[0202] 9: store total-array[] to the task node

[0203] 10: end while

[0204] 11: end if

[0205] In summary, the alternatives, although theoretically supported, are generally inferior to the present technical solution in terms of deployment complexity, execution cost, and expansion flexibility. Considering the runtime efficiency, test scale expansion capability, and early adaptation cost required for fuzz testing, the present application has the advantages of simple structure, excellent performance, and flexible synchronization granularity, making it the most valuable solution for current engineering applications.

[0206] In summary, the present application provides a parallel fuzz testing node optimization method for distributed environments, which has the following beneficial effects:

[0207] Firstly, a polling path synchronization mechanism is proposed and implemented in the actual distributed fuzzing system. This mechanism bypasses the architecture burden of traditional master nodes actively pushing coverage paths, and adopts the strategy of distributed test nodes self-maintaining "jump array". All path data is abstracted into a simple one-dimensional bit array, which is collected by the master node in a one-way polling manner. Through simple bitwise or merging logic, information synchronization is achieved. This technology has strong scalability and can still guarantee synchronization efficiency when the number of test nodes is large, avoiding a significant increase in the probability of errors.

[0208] Secondly, the jump number interval precision matching algorithm is introduced in the synchronization process. It not only compares the path coverage at the bit level, but also records the frequency of basic block appearance on the jump path. By setting jump trigger segmented intervals such as [1-10], [10-100], [100+], the granularity of path scarcity analysis is improved, preventing the "pseudo-coverage" problem of covering paths but not activating key behaviors. This processing enables the synchronization mechanism to change from "macroscopic path to local state".

[0209] Thirdly, based on the delayed lazy synchronization strategy, the master node does not immediately dispatch instructions after each synchronization, but through shared memory mapping (or RPC) method, the test node loads new configuration when it knows that the path state meets the characteristic value change condition, thus completely avoiding blocking waiting and better adapting to the asynchronous fuzzer life cycle.

[0210] Fourthly, a queue fusion model based on seed quality weighted scoring is constructed for the processing of the fuzzy task queue of the test node. The evaluation indexes are strictly proportionally weighted and dynamically balanced. When a queue has too many trivial seeds, the weight of high-quality seeds is automatically increased to alleviate the queue convolution trend and inject test vitality into the low-efficiency node. Specifically, in the implementation of the queue fusion mechanism, when it is detected that the proportion of trivial seeds (i.e., seeds that have not covered new jumps for many rounds in succession, have too high genetic affinity, and have low jump danger score) in the fuzzy task queue of a test node exceeds a dynamic threshold (e.g., 60%), the weight adjustment mechanism is triggered. The mechanism dynamically optimizes the weight distribution of the four evaluation indexes: for example, the upper limit of the weight of the jump coverage rate index is increased to 60%, the upper limit of the weight of the jump danger index is increased to 30%, the lower limit of the weight of the genetic affinity index is suppressed to 15%, and the weight of the sample data size index remains unchanged at 10%. After the weight adjustment, the comprehensive score of all seeds in the queue of the node is recalculated based on the four evaluation indexes, so that the relative value of high-value seeds (e.g., seeds that cover a large number of new jumps or trigger high-risk behaviors) that are originally submerged due to low absolute scores is significantly improved. These high-quality seeds with score jumps will be preferentially migrated to low-efficiency nodes in the subsequent queue merging process, breaking the test stalemate caused by the excessive proliferation of trivial seeds and effectively activating the global exploration vitality. The entire adjustment process is implemented through the weight coefficient table maintained by the master node, and a weight change rate limit of, for example, not more than 20% per minute is set to ensure smooth transition of the system.

[0211] Fifthly, in order to be compatible with the fuzzer configuration module, a seed affinity tracking ID pool system is further designed, each test sample is assigned with its source ID and affinity depth information, and depth backtracking is used to avoid "seed polling trap", i.e., the problem of concentrated path coverage caused by the high-speed derivation of homologous seeds due to the common fusion of multiple nodes.

[0212] In summary, the present application proposes a new architecture of a distributed fuzzy test synchronization and optimization framework, which systematically solves key problems such as serious redundancy and repetition of fuzzy tests in a multi-node environment and unbalanced seed quality from the bottom path bitmap structure modeling, polling synchronization mechanism control to task queue state allocation mechanism. The actual test shows that compared with the existing distributed system based on the state bitmap polling mechanism and the dynamic queue fusion strategy, the fuzzy test platform can improve the path discovery speed by 52.4% and reduce the high repetition rate event by 73.1%, significantly improving the parallel efficiency, vulnerability discovery rate and node load coordination degree of the fuzzy test, and has wide technical promotion and commercial application value.

[0213] In a second aspect, the application provides a parallel fuzzing test node optimization system for a distributed environment, applied to the parallel fuzzing test node optimization method for a distributed environment. The multiple test nodes include at least a first test node and a second test node. The system includes:

[0214] An initial unit is configured to perform, asynchronously, fuzz test generation on the first test node by using a seed sample of a fuzz task queue, and record branch jump addresses and execution times reached during the test process to obtain a local jump path array corresponding to the first test node.

[0215] A polling unit is configured to perform polling collection on the local jump path arrays of the multiple test nodes by using a master node, and perform a bitwise OR operation to obtain a global path array.

[0216] A matching unit is configured to perform matching of jump addresses and execution times between the global path array and a local jump path array of the second test node. If the matching result is repeated, the mutation and test process of the current seed sample are stopped, and the next seed sample is selected from the local fuzz task queue to test the second test node to obtain a local jump path array corresponding to the second test node.

[0217] An evaluation unit is configured to periodically evaluate the value of the seed sample of the fuzz task queue of the test node to obtain a score of the seed sample.

[0218] A result unit is configured to select, in at least one period, two test nodes with the highest and lowest scores, and select some seed samples with high scores from the test node with the highest score to merge into the test node with the lowest score, update the fuzz task queue of the test node with the lowest score, and obtain the optimization of the test node.

[0219] In a third aspect, the application provides a computer device including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the parallel fuzzing test node optimization method for a distributed environment are implemented.

[0220] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0221] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments.

[0222] The scope of protection of the present disclosure is not limited to the above-mentioned embodiments. Obviously, those skilled in the art can make various modifications and changes to the present disclosure without departing from the scope and spirit of the present disclosure. If these modifications and changes belong to the scope of the claims of the present disclosure and its equivalent technologies, the present disclosure also includes these modifications and changes.

Claims

1. A parallel fuzzing node optimization method for a distributed environment, a plurality of test nodes comprising at least a first test node and a second test node, characterized in that, The method comprises: a seed sample of a fuzzy task queue is used to execute test cases generated by fuzzy testing on the first test node asynchronously, and branch jump addresses and execution times reached in the test process are recorded to obtain a local jump path array corresponding to the first test node; a master node is used to poll and collect local jump path arrays of multiple test nodes, and perform bitwise OR operation to combine to obtain a global path array; the global path array and a local jump path array of the second test node are used to match jump addresses and execution times, if the matching result is repeated, the mutation and test process of the current seed sample are stopped, the next seed sample is selected from the local fuzzy task queue to test the second test node, and a local jump path array corresponding to the second test node is obtained; the seed sample of the fuzzy task queue of the test node is periodically evaluated to obtain a score of the seed sample; in at least one period, two test nodes with the highest and lowest scores are screened, and part of the seed samples with high scores in the test node with the highest score are selected and combined into the test node with the lowest score, the fuzzy task queue of the test node with the lowest score is updated, and the test node is optimized.

2. The parallel fuzzing node optimization method for distributed environment according to claim 1, wherein, The step of using a seed sample of a fuzzy task queue to execute test cases generated by fuzzy testing on the first test node asynchronously, and recording branch jump addresses and execution times reached in the test process to obtain a local jump path array corresponding to the first test node comprises: a seed sample of a fuzzy task queue is pre-stored in the first test node to obtain test cases generated by fuzzy testing; the test cases are executed asynchronously in the first test node, and branch jump addresses and execution times are recorded to obtain a local jump path array corresponding to the first test node.

3. The parallel fuzzing node optimization method for distributed environment of claim 1, wherein, The step of using a master node to poll and collect local jump path arrays of multiple test nodes, and performing bitwise OR operation to combine to obtain a global path array comprises: the master node sends data requests to multiple test nodes according to a preset order to poll and collect local jump path arrays of the multiple test nodes; the master node performs bitwise OR operation to combine the global path array and the local jump path arrays of the multiple test nodes to obtain the global path array of the master node.

4. The parallel fuzzing node optimization method for distributed environment of claim 1, wherein, The step of using the global path array and a local jump path array of the second test node to match jump addresses and execution times, if the matching result is repeated, stopping the mutation and test process of the current seed sample, and selecting the next seed sample from the local fuzzy task queue to test the second test node to obtain a local jump path array corresponding to the second test node comprises: the global path array and the local jump path array of the second test node are used to match jump addresses to obtain a jump address matching result; if the jump address matching result is repeated, execution times are matched to obtain an execution time matching result; If the execution frequency matching result is repeated, the mutation and test process of the current seed sample of the second test node are stopped, the next seed sample is selected from the local fuzzy task queue to test the second test node, and a local jump path array corresponding to the second test node is obtained.

5. The parallel fuzzing node optimization method for distributed environment according to claim 4, wherein, The step of performing a bitwise OR operation on the global path array of the master node and the local jump path array of each test node to obtain the global path array of the master node comprises: The first test node is polled and collected by using the global path array of the master node, the local jump path array of the first test node is merged into the global path array, and a new global path array is obtained. The new global path array is sent to the second test node, and the second test node updates its global path array copy.

6. The parallel fuzzing node optimization method for distributed environment of claim 1, wherein, The step of periodically evaluating the value of the seed sample in the fuzzy task queue of the test node to obtain the score of the seed sample comprises: Construct a seed sample value score index and assign a corresponding weight, wherein the seed sample value score index comprises jump coverage, genetic kinship, jump risk, and sample data size; Periodically, the value of the seed sample in the fuzzy task queue of the test node is evaluated according to the seed sample value score index, and the score of the seed sample is obtained. According to the score of the seed sample, the seed sample is sorted.

7. The parallel fuzzing node optimization method for distributed environment of claim 1, wherein, The step of screening the two test nodes with the highest and lowest scores in at least one period, and selecting part of the high-score seed samples in the test node with the highest score to merge into the test node with the lowest score to update the fuzzy task queue of the test node with the lowest score to obtain the optimization of the test node comprises: In at least one period, the score of all seed samples of each test node is obtained to obtain the score of the fuzzy task queue; According to the score of the fuzzy task queue, the score of each test node is determined; According to the score of each test node, the two test nodes with the highest and lowest scores are obtained. In the test node with the highest score, a preset proportion of high-score seed samples are selected to merge into the test node with the lowest score to obtain the optimization of the test node.

8. The parallel fuzzing node optimization method for distributed environment of claim 1, wherein, The step of using the seed sample of the fuzzy task queue to asynchronously execute the test case generated by the fuzzy test on the first test node, and recording the branch jump address and execution frequency reached during the test process to obtain the local jump path array corresponding to the first test node comprises: A unique ID and kinship depth information are constructed for each seed sample; When the seed sample is mutated to form a mutated seed sample, the mutated seed sample inherits the parent seed ID and the kinship depth is increased.

9. A parallel fuzzing node optimization system for a distributed environment, characterized by, The parallel fuzzy test node optimization method for distributed environment according to any one of claims 1-8, wherein the plurality of test nodes at least comprises a first test node and a second test node, and the system comprises: An initial unit is configured to perform, asynchronously, a fuzzy test on the first test node by using a seed sample of a fuzzy task queue, and record a branch jump address and an execution frequency reached during the test to obtain a local jump path array corresponding to the first test node; A polling unit is configured to collect, by using a master node, the local jump path arrays of the multiple test nodes, and perform a bitwise OR operation to obtain a global path array; A matching unit is configured to match the global path array with a local jump path array of the second test node in terms of the branch jump address and the execution frequency, and if the matching result is repeated, stop mutation and test of the current seed sample, and select a next seed sample from the local fuzzy task queue to test the second test node to obtain a local jump path array corresponding to the second test node; An evaluation unit is configured to periodically evaluate the seed sample of the fuzzy task queue of the test node to obtain a score of the seed sample; A result unit is configured to select, in at least one period, two test nodes with the highest and lowest scores, and select some seed samples with high scores from the test node with the highest score to merge into the test node with the lowest score, update the fuzzy task queue of the test node with the lowest score, and obtain an optimization of the test node. 10.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-9. The processor executes the computer program to implement the steps of the parallel fuzzy test node optimization method for a distributed environment according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and apparatus for selecting fuzzy test case

    CN109062795A

  • Test data construction method and equipment

    CN118672893A

  • Fuzzy testing a software system

    US20230367704A1