Protocol fuzz testing method and system based on fine-grained state division and selection
Through the two-layer state recognition mechanism and optimized state selection strategy, the problems of excessive granularity of state division and insufficient state selection algorithm in the prior art are solved, and efficient and accurate protocol fuzz testing is achieved.
Patent Information
- Application Number
- CN202510145369.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The state division granularity in the prior art in protocol fuzz testing is too coarse, resulting in state confusion and unnecessary consumption of test resources; at the same time, the state selection algorithm fails to fully consider the topological characteristics and functional roles of the state nodes, which affects the reasonable allocation of test resources.
A two-layer state recognition mechanism is adopted to build basic state identification through the combination of response code and control fields, and state subdivision is performed using local sensitive hash algorithm. Combined with the optimized state selection strategy, analyze the topological structure and node characteristics of the global state transition graph, and conduct comprehensive evaluation and selection.
It significantly improves the efficiency and accuracy of protocol fuzz testing, avoids state confusion and redundancy, realizes key tests of key control nodes, and improves the rational allocation and utilization of test resources.
Smart Images

Figure CN119996271A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fuzzy testing, and in particular relates to a protocol fuzzy testing method and system based on fine-grained state division and selection. Background Art
[0002] As a set of rules agreed upon by both parties in network communication, network protocols specify the format, method and error verification mechanism of data transmission to ensure the reliable exchange of data between different devices and systems. Although these protocols are designed to ensure the security and consistency of communication, any complex software implementation may introduce vulnerabilities. And with the further iteration of the protocol, more complex branches are introduced, increasing the possibility of vulnerabilities. In recent years, the field of protocol fuzz testing has made significant progress, and the research focus has mainly focused on stateful gray-box fuzz testing methods. These methods greatly improve the test efficiency and exploration of the state space by introducing state-aware mechanisms to guide the generation of test cases, and promote fuzz testing towards more accurate state exploration.
[0003] Chinese patent CN117478566A proposes an intelligent fuzzy testing scheme for text protocols. The scheme implements field segmentation by identifying the separation position in the message, establishes a distribution feature library of field values, and adjusts the application intensity of the mutation strategy according to the richness and rarity of the field values. This field-aware method improves the pertinence of test case generation. However, the state representation method used in this method lacks in-depth analysis of the internal execution characteristics of the protocol, which makes it difficult for the method to accurately identify and distinguish the actual protocol state, and easily causes incorrect state division. At the same time, the characteristics and functions of the state nodes are not fully considered in the state selection strategy, which affects the reasonable allocation of test resources.
[0004] Chinese patent CN118784340A proposes a state monitoring mechanism based on multi-dimensional indicators. This method comprehensively considers the test frequency, number of reachable paths and test effect of the state, and balances the state selection strategy through these indicators to achieve balanced testing of paths of different depths. This method has made some progress in optimizing state selection, but its definition of state still adopts a more traditional approach and fails to deeply analyze and subdivide the internal state characteristics of the protocol. This coarse-grained state representation method limits the accuracy of the test, and even with a good state selection strategy, it is difficult to fully realize its potential in vulnerability mining. This limitation illustrates the necessity of a more detailed division of protocol states.
[0005] In summary, in the field of fuzz testing of stateful protocols, the existing technologies have problems in state perception and selection and scheduling of state nodes:
[0006] The state division granularity is too coarse. For state division, the existing technology mainly relies on the response code of the protocol to define the state. This simple state representation method cannot accurately reflect the internal state characteristics of the protocol, which leads to two problems: on the one hand, since the correspondence between the response code and the actual state is not one-to-one, the system may mistakenly divide states that are essentially the same but have different response codes into different categories, resulting in unnecessary consumption of test resources; on the other hand, different internal states may produce the same response code. This state confusion reduces the pertinence and effectiveness of the test.
[0007] The state selection algorithm has defects. In terms of state selection, although the existing methods have proposed a variety of state evaluation indicators, they have failed to fully consider the topological characteristics and functional roles of the state nodes in the protocol state machine. The limitations of this selection strategy make it impossible to optimize the allocation of test resources according to the actual importance of the nodes, thus affecting the overall efficiency of the test. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a protocol fuzzy testing method and system based on fine-grained state division and selection, which effectively improves the efficiency and accuracy of protocol fuzzy testing by designing a double-layer state recognition mechanism and combining it with an optimized state selection strategy.
[0009] The present invention provides a protocol fuzzy testing method based on fine-grained state division and selection, comprising the following steps:
[0010] S1: Fuzz test preprocessing stage: deploy the entity program of the protocol to be tested and configure the initial seed set;
[0011] S2: Basic state identification stage; receive complete feedback information of each initial seed in the initial seed set, including the final coverage edge information and the response message sequence returned by the protocol entity program to be tested, establish the correspondence between the sent message and the response message, then extract the response code from the response message, and construct a joint state identification in combination with the control field of the corresponding sent message, and use the information to build and maintain the global state transition graph; and use all the initial seeds in the initial seed set as test seeds;
[0012] S3: Double-layer state segmentation stage; select the seeds saved in each state, form a corresponding state seed set and replay the relevant message sequence to guide the server to the target state, then send a specially constructed detection message and collect the response code and execution path feedback information generated in the test; preliminarily classify the seeds according to the response code, and then use the local sensitive hashing algorithm to analyze the path characteristics of each seed; by comparing the similarity of the path increments, the seeds with similar path characteristics are classified into the same refined state;
[0013] S4: State evaluation and selection stage: conduct a comprehensive evaluation of the state nodes, systematically analyze the topological structure of the global state transition graph to identify error state nodes and key control nodes, and conduct a comprehensive analysis based on the node's in-degree characteristics, transition probability and improved betweenness centrality index;
[0014] S5: Fuzz test execution phase; select the target state and seed according to the evaluation result of step S4, and perform random mutation on them to continuously generate new seeds; use the mutated new seeds to execute the entity program of the protocol to be tested, and collect the coverage information and response message sequence during its execution process, and then use the analysis mechanism established in steps S2 and S3, including state identification extraction and path feature analysis, to characterize the state characteristics of the new execution results;
[0015] S6: Feedback analysis phase; systematically analyze the test execution results and maintain global status information.
[0016] Preferably, the specific steps of step S1 include:
[0017] S1.1: Deploy the complete file of the protocol entity program to be tested, configure the environment and dependencies required by the protocol entity program to be tested, and compile it using the afl-gcc stub compilation tool;
[0018] S1.1.1: Correctly deploy the complete file of the protocol entity program to be tested in the local environment computer, and configure the environment and dependencies required for its operation; use the compilation tool to pre-compile the protocol entity program to be tested, and specifically identify and distinguish each basic block in the code during the compilation process; a basic block refers to a linearly executed instruction sequence in the program, which has the characteristics of a single entry point and a single exit point;
[0019] S1.1.2: Use the instrumentation compilation tool to instrument the protocol entity program to be tested, insert monitoring code at the entrance of each basic code block, and assign a unique random identification value to each basic block; when the program execution transfers from one basic block to another, the random identification values of the two basic blocks before and after are combined into an edge identifier; by recording and counting these edge identifiers, the code paths covered during program execution can be accurately tracked, providing the necessary coverage feedback for subsequent analysis, where "edge" and "coverage edge" specifically refer to the jump behavior of the basic blocks monitored when the protocol entity program to be tested is running;
[0020] S1.2: Prepare the initial seed library; locally start the executable file of the entity program of the protocol to be tested, and perform multiple rounds of normal communication with it by simulating the client locally; during the communication process, use the network packet analysis tool to capture the communication traffic and extract the message data sent by the client as the initial seed; start the fuzz test with at least one legal initial seed;
[0021] S1.2.1: Start a network packet analysis tool locally and continuously monitor the local network interface to capture communication traffic;
[0022] S1.2.2: Start a protocol entity program to be tested locally as a server, and start a client program separately; simulate the normal communication process for testing according to the RFC specification document; during this process, the network packet analysis tool captures the message sequence sent by the client, combines the captured multiple messages in chronological order and splits them with specific delimiters to form a usable initial seed;
[0023] S1.2.3: As needed, repeat step S1.2.2 to obtain initial seeds in multiple different scenarios to form an initial seed set.
[0024] Preferably, the specific steps of step S2 include:
[0025] S2.1: Receive the complete feedback information of each initial seed in step S1, including the final coverage edge information and the response message sequence returned by the protocol entity program under test; identify the response code of each response message and extract the control field information from the corresponding sent message;
[0026] S2.2: Construct a state identifier and update the state transition information; construct a state identifier based on the extracted response code and control field information, and use them to construct and maintain relevant information of the global state transition graph;
[0027] S2.3: Determine the seed persistence strategy; persist or clear the current test seed according to its execution effect. For the test seed that needs to be saved, add it to the persistent seed library and associate it with the seed set of all nodes on its state transfer path.
[0028] Preferably, the specific steps of step S2.1 include:
[0029] S2.1.1: Establish a correspondence between sent messages and response messages; each sent message corresponds to a response message;
[0030] S2.1.2: Get the response code of each response message. Execute the callback function GetCode for each response message. The callback function GetCode extracts and processes certain fields from any response message, and finally hashes them into an integer variable, and regards the integer variable as the response code of the response message;
[0031] S2.1.3: Take out the control field of each sent message; execute the callback function GetControl for each message in the test seed. The callback function GetControl extracts certain fields from any sent message and processes them, and finally returns a character variable, and regards the character variable as the control field of the sent message.
[0032] Preferably, the specific steps of step S2.2 include:
[0033] S2.2.1: Construct a state joint identifier; for each set of sent messages and corresponding response messages, construct a two-tuple 〈Control_Byte, Response_code〉 as a state identifier, where Control_Byte is the control field of the sent message, and Response_code is the response code of the response message;
[0034] S2.2.2: Define and construct a global state transition graph. The system maintains a global directed graph structure to represent the state space of the protocol entity program to be tested; each state node v is uniquely identified by a tuple 〈Control_Byte,Response_code〉, where Control_Byte is the message control field that triggers the state, and Response_code is the corresponding response code; the state node also records the cumulative number of times it has been accessed, and maintains a seed set that can reach the state; and node 0 represents the initial state of the protocol entity program to be tested;
[0035] S2.2.3: Maintain the state transition relationship; perform timing processing on the message sequence in the test seed; for each group of sent messages and response messages, construct its state identification tuple 〈Control_Byte, Response_code〉; therefore, one test seed corresponds to several state identification tuples 〈Control_Byte, Response_code〉; check each state identification tuple 〈Control_Byte, Response_code〉, and if it has not been recorded, create a corresponding state node in the global state transition graph; then the system establishes directed edges between each pair of adjacent state nodes, and updates the statistical information such as the node's access count, seed set, and transition frequency.
[0036] Preferably, the specific steps of step S2.3 include:
[0037] S2.3.1: Determine whether the test seed needs to be saved; check whether a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time during the execution of the test seed. When a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time, the test seed is marked as a state to be saved;
[0038] S2.3.2: Execute seed persistence; for the marked test seeds, add their complete contents to the persistent seed library for storage; at the same time, associate the test seed with all state nodes on its state transfer path, ensuring that the seed set of each relevant node contains the test seed. Through this association mechanism, a dedicated seed set is established and maintained for each state node. The seeds in the seed set have the ability to set the state of the entity program of the protocol to be tested to the corresponding state by replaying part of the message sequence.
[0039] Preferably, the specific steps of step S3.1 include:
[0040] S3.1: Execute state replay and detection; for the state nodes that need to be subdivided, replay the seed set for testing; replay the messages in the seed until the state of the protocol entity program to be tested is set to the current state, calculate the coverage edge information of the current execution, and then send a specific detection message to collect new execution coverage edge information, and then calculate the path increment feature and response code;
[0041] S3.1.1: Select a subdivision state node; periodically monitor the global state node, and when the number of seeds that have not been subdivided in a state node exceeds a preset threshold of 50 seeds, mark it as a state to be subdivided;
[0042] S3.1.2: Replay seeds: For each seed in the state to be subdivided, replay its message sequence until the protocol entity program to be tested is guided to the current state; record the complete coverage edge information at this time as the reference path P_pre, and then send a message of the same special structure to all seeds, at this time obtain the current complete coverage edge information P_cur and the current response code information, and then use the reference path P_pre and the complete coverage edge information P_cur to calculate the incremental path P_m;
[0043] S3.1.3: Calculate the path increment; analyze the difference between the reference path P_pre and the complete coverage edge information P_cur, and calculate the incremental path P_m;
[0044] Create an empty bitmap array, and then compare each byte in the reference path P_pre and the complete coverage edge information P_cur one by one; when the i-th byte is found to be different, the system sets the (i&7)th bit of the (i>>3)th byte in the bitmap to 1, and obtains the incremental path P_m;
[0045] S3.2: Perform the first level division based on the response code; analyze the response codes of all seeds to the same special detection message, and classify the seeds with the same response code into the same sub-state, thereby completing the first level of state subdivision;
[0046] S3.3: Perform the second-level partitioning based on the incremental path. For the state seed set in each sub-state, the second-level partitioning is performed using the local sensitive hashing algorithm according to the incremental path P_m calculated in S3.1.
[0047] Preferably, the specific steps of step S4 include:
[0048] S4.1: Global state evaluation and selection; Analyze the characteristics of each node in the global state transition graph and evaluate and score it; Before each round of fuzz testing, comprehensively score the node based on its topological characteristics, historical benefits, and state characteristics, and select the appropriate state for the next round of testing;
[0049] S4.1.1: Simple state selection; when the number of fuzz tests is less than 10,000, a simple state selection strategy is used, that is, polling and selecting each node in the current global state graph to ensure that the state space of the entity program of the protocol to be tested is fully explored in the early stage of fuzz testing; otherwise, go to step S4.1.2;
[0050] S4.1.2: Calculate time decay benefits. For each state node, introduce a time decay mechanism to evaluate its historical test benefits. This mechanism gives recent test findings a higher evaluation value by giving historical findings a weight that decreases over time. The specific calculation formula is:
[0051]
[0052] Among them, γ takes the value of 0.95 as the time decay coefficient, n is the current test round, t is the historical round index, paths_discovered(v,t) represents the number of new seeds discovered by state node v in the tth test round, and selected_times(v) represents the total number of times state node v is selected for fuzz testing;
[0053] S4.1.3: Identify error states; count the in-degree of each state node, calculate the number of transition edges from all other state nodes to the node, and divide it by the total number of state nodes to get the in-degree ratio; when the in-degree ratio of a state node exceeds 0.7, the node is judged as a potential error state, and its score weight is reduced to 0.3 times, thereby reducing the tendency to test such states;
[0054] S4.1.4: Analyze the transition probability; calculate the transition frequency between all state nodes, count all possible transition targets of state node v and the corresponding number of transitions, and then calculate its transition probability score; the specific calculation formula is:
[0055]
[0056] Among them, V is the set of state nodes, trans(i,v) represents the number of historical transfers from state node i to state node v, Represents the total number of transitions to all sub-state nodes to which state node i can transfer;
[0057] S4.1.5: Evaluate node betweenness centrality; introduce betweenness centrality calculation to evaluate the importance of a state node in the global state transition graph. Update the betweenness centrality score of the node every 10 rounds of state selection. Quantify its importance by calculating the frequency of the state node appearing in the shortest path from the initial state to other states. The specific calculation formula is:
[0058]
[0059] Among them, σst(v) represents the number of paths containing node v in the shortest path from the initial state 0 to any other state, and σst represents the total number of shortest paths from the initial state to other states, that is, the total number of state nodes;
[0060] S4.1.6: Final score calculation: Taking into account the evaluation results of the above dimensions, the final score of each state node is calculated. The specific calculation formula is:
[0061] Score(v)=R(v)·α(v)·(1+BC(v))+Score trans (v)
[0062] Where R(v) is the time-decayed historical return value calculated in the previous step. The coefficient α(v) takes the value of 0.3 only when the in-degree ratio of a node exceeds 0.7, otherwise it takes the value of 1.0. BC(v) represents the improved betweenness centrality score of the node, which is used to reflect the importance of the node in state transition. Score_trans(v) is the transition probability score of the node.
[0063] After obtaining the original scores of all nodes, normalization is performed: the sum of all node scores is calculated, and then the score of each node is divided by the sum to obtain a normalized probability distribution; based on this probability distribution, a roulette wheel selection algorithm is used to determine the target state for the next round of testing, so that states with higher scores have a greater probability of being selected, while ensuring that states with lower scores still have a chance to be explored;
[0064] S4.2: Specific subdivision state seed selection; check whether there are new seeds that have not been fuzz tested in the selected state, if so, give priority to these new seeds for testing; if there are no new seeds, the system randomly selects a subdivision state in the state based on uniform probability distribution, and then randomly selects a seed in the selected subdivision state as the test object; after completing the seed selection, determine the specific message for the fuzz test; analyze the complete state transition path of the seed, and determine the message sequence required to guide the protocol entity program to be tested to the target state, and then select the next message corresponding to the state transition point in the sequence as the fuzz test object; pass the selected seed and the specific message location information that needs to be fuzz tested to step S1 as the input parameter for a new round of fuzz testing.
[0065] Preferably, the specific steps of step S5 include:
[0066] S5.1: Enter the formal fuzz testing phase; in each round of testing, select the target state and its corresponding seed for testing based on the evaluation results of step S4;
[0067] S5.2: Generate test seeds for the currently selected seeds using a mutation strategy based on a genetic algorithm;
[0068] S5.2.1: Determine the fuzzy area; since a single seed is composed of multiple messages, this stage first obtains the target seed to be tested and the specific message position that needs to be fuzz tested through the interface of step S4; in the subsequent mutation process, the specific message is mutated and the mutated message is replaced back to the original position, thereby constructing a new test seed;
[0069] S5.2.2: Enter the deterministic mutation stage; perform mutation operations on each byte position of the target message in sequence according to a preset fixed order, including the steps of bit-by-bit flipping, byte-by-byte flipping, and byte-by-byte replacement of special values;
[0070] S5.2.3: Enter the non-deterministic mutation stage; randomly select and combine multiple mutation operators to mutate the target message. The mutation operators include random position insertion, random block deletion, and block splicing and reorganization. In each round of mutation, the operator combination and mutation position used are randomly determined to generate a new test message.
[0071] S5.3: Start the protocol entity program to be tested, send the mutated test seed for testing, collect the execution results, including program coverage data and the complete response message returned by the protocol entity program to be tested, and pass the feedback information to step S2 through the interface for analysis;
[0072] S5.3.1: Start the protocol entity program to be tested. The test seed to be sent consists of the mutated target message and other original messages that have not been mutated in sequence. Establish a network connection with the protocol entity program to be tested and send the messages one by one according to the timing of the messages. For each message, after the message is sent, periodically check the edge coverage feedback of the protocol entity program to be tested. When the edge coverage information of two consecutive checks is consistent, it is determined that the message has been completely processed, and then the next message in the sequence is sent until all the messages in the test seed are sent.
[0073] S5.3.2: When all messages in the test seed have been sent and responses received, the system organizes the test execution results, including the final coverage edge information and the complete response message sequence, and passes this feedback information to step S2 through the interface for analysis;
[0074] S5.4: Repeat steps S5.2 to S5.3 until the preset upper limit of mutation times is reached; then return to step S5.1 to select a new state and seed for the next round of fuzzy testing.
[0075] A system for protocol fuzz testing method based on fine-grained state division and selection, comprising a state division module, a state evaluation module, a test execution module and a feedback analysis module; each module works together to achieve efficient fuzz testing of the protocol entity program to be tested;
[0076] Fuzz test preprocessing module; responsible for deploying the entity program of the protocol to be tested;
[0077] State segmentation module: This module is responsible for implementing a two-layer state recognition mechanism, including a basic feature extraction unit and a state segmentation processing unit. The basic feature extraction unit is responsible for extracting the response code and control field information from the interaction process and constructing the initial state identifier. The state segmentation processing unit works regularly to further segment the state through path analysis and similarity clustering.
[0078] State evaluation module: This module is responsible for systematically analyzing and evaluating the global state transition graph and guiding the allocation of fuzz testing resources. In the early stage of fuzz testing, the module selects state nodes by random polling to ensure full exploration of the initial state space. When enough test samples are accumulated, the module switches to the heuristic evaluation mechanism. By analyzing the topological structure of the state transition graph, the module identifies the error state nodes and key control nodes, and conducts a comprehensive analysis based on the node's in-degree characteristics, transition probability and improved betweenness centrality index. In the scoring process, the module introduces a time decay mechanism to calculate the node's historical benefits, and combines it with the node feature score, and finally generates a score for each state node that reflects its importance. Through this mechanism that combines random exploration and heuristic evaluation, the module provides a reliable decision-making basis for the reasonable allocation of test resources.
[0079] Test execution module: This module is responsible for executing the specific fuzz testing process. Based on the target state and seed provided by the state evaluation module, this module performs mutation testing on the target message specified therein. During the mutation process, a new test seed is generated by randomly combining multiple mutation operators while keeping other messages in the test seed unchanged. For each mutated test seed, the module starts the entity program of the protocol to be tested and sends a message sequence in sequence.
[0080] Feedback analysis module: This module is responsible for systematically analyzing the test execution results and maintaining the global state information; this module collects the program execution path and response characteristics of each test, analyzes the state transition sequence and updates the statistical data of node access frequency and transition probability in the global state transition graph.
[0081] The present invention has the following technical effects:
[0082] 1. The efficiency and accuracy of protocol fuzz testing are significantly improved through the two-layer state identification and dynamic evaluation mechanism. In terms of state identification, the system first builds a basic state identifier based on the combination of response code and control field, and then implements state segmentation through path feature analysis, effectively avoiding state confusion and redundancy. In terms of state evaluation, the system implements focused testing of key control nodes by analyzing the topological characteristics, historical benefits and transfer probabilities of nodes.
[0083] 2. The present invention realizes full-process automated testing. Through the collaborative work of multiple modules, the system can automatically complete the entire process from state identification, evaluation selection to test execution, greatly reducing the testing cost and the need for manual intervention. At the same time, the system uses a local sensitive hashing algorithm to perform state segmentation, quickly identify similar paths, and avoid the performance overhead of directly calculating path similarity. Through an efficient automation mechanism, the present invention can maintain stable testing efficiency in a low-speed network protocol fuzzy testing environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 It is a schematic diagram of the overall framework of the system of the present invention;
[0085] Figure 2 This is a diagram of the state subdivision method in the double-layer state subdivision stage. DETAILED DESCRIPTION
[0086] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings.
[0087] like Figure 1 As shown, the protocol fuzz testing method based on fine-grained state division and selection includes the following steps:
[0088] S1: Fuzz test preprocessing stage: deploy the entity program of the protocol to be tested and configure the initial seed set;
[0089] S1.1: Deploy the complete file of the protocol entity program to be tested, configure the environment and dependencies required by the protocol entity program to be tested, and compile it using the afl-gcc stub compilation tool;
[0090] S1.1.1: Correctly deploy the complete file of the protocol entity program to be tested in the local environment computer, and configure the environment and dependencies required for its operation; use the compilation tool to pre-compile the protocol entity program to be tested, and specifically identify and distinguish each basic block in the code during the compilation process; a basic block refers to a linearly executed instruction sequence in the program, which has the characteristics of a single entry point and a single exit point;
[0091] S1.1.2: Use the instrumentation compilation tool to instrument the protocol entity program to be tested, insert monitoring code at the entrance of each basic code block, and assign a unique random identification value to each basic block; when the program execution transfers from one basic block to another, the random identification values of the two basic blocks before and after are combined into an edge identifier; by recording and counting these edge identifiers, the code paths covered during program execution can be accurately tracked, providing the necessary coverage feedback for subsequent analysis, where "edge" and "coverage edge" specifically refer to the jump behavior of the basic blocks monitored when the protocol entity program to be tested is running;
[0092] S1.2: Prepare the initial seed library; locally start the executable file of the entity program of the protocol to be tested, and perform multiple rounds of normal communication with it by simulating the client locally; during the communication process, use the network packet analysis tool to capture the communication traffic and extract the message data sent by the client as the initial seed; start the fuzz test with at least one legal initial seed;
[0093] S1.2.1: Start a network packet analysis tool locally and continuously monitor the local network interface to capture communication traffic;
[0094] S1.2.2: Start a protocol entity program to be tested locally as a server, and start a client program separately; simulate the normal communication process for testing according to the RFC specification document; during this process, the network packet analysis tool captures the message sequence sent by the client, combines the captured multiple messages in chronological order and splits them with specific delimiters to form a usable initial seed;
[0095] S1.2.3: Repeat step S1.2.2 as needed to obtain initial seeds in multiple different scenarios to form an initial seed set. Although fuzz testing requires at least one valid seed to start, a diverse initial seed set can provide wider test coverage.
[0096] S2: Basic state identification stage; receive complete feedback information of each initial seed in the initial seed set, including the final coverage edge information and the response message sequence returned by the protocol entity program to be tested, establish the correspondence between the sent message and the response message, then extract the response code from the response message, and construct a joint state identification in combination with the control field of the corresponding sent message, and use the information to build and maintain the global state transition graph; and use all the initial seeds in the initial seed set as test seeds;
[0097] S2.1: Receive the complete feedback information of each initial seed in step S1, including the final coverage edge information and the response message sequence returned by the protocol entity program under test; identify the response code of each response message and extract the control field information from the corresponding sent message;
[0098] S2.1.1: Establish a corresponding relationship between the sent message and the response message; each sent message corresponds to a response message; the system adopts a one-to-one mapping method, that is, each sent message corresponds to a response message.
[0099] S2.1.2: Get the response code of each response message. The callback function GetCode is executed for each response message. The callback function GetCode extracts and processes certain fields from any response message, and finally hashes them into an integer variable, and regards the integer variable as the response code of the response message; for example, in the test of the RTSP protocol, a class of implementations returns "200OK....." or "405Method ERROR....." At this time, the implementation of the callback function GetCode takes out the first integer of the message, that is, 200 and 405, as the response code.
[0100] S2.1.3: Take out the control field of each sent message; execute the callback function GetControl for each message in the seed. The callback function GetControl extracts certain fields from any sent message and processes them, and finally returns a character variable, and regards the character variable as the control field of the sent message. Taking the RTSP protocol as an example, for a message in the form of "OPTIONS rtsp: / 127.0.0.1:8554 / mystream RTSP / 1.0CSeq....", the callback function GetControl extracts the first field "OPTIONS" as the control field. In addition, in order to prevent the expansion of the state space caused by mutation, the callback function GetControl must pre-define a legal control field set. In the RTSP protocol, legal control fields include: OPTIONS, DESCRIBE, SETUP, PLAY, PAUSE, TEARDOWN, ANNOUNCE, RECORD, REDIRECT, and SET_PARAMETER. When the system encounters a control field that is not in this set, it will be uniformly marked as "UnKnow" to avoid unnecessary expansion of the state space;
[0101] S2.2: Construct a state identifier and update the state transition information; construct a state identifier based on the extracted response code and control field information, and use them to construct and maintain relevant information of the global state transition graph;
[0102] S2.2.1: Construct a state joint identifier; for each set of sent messages and corresponding response messages, construct a tuple 〈Control_Byte,Response_code〉 as a state identifier, where Control_Byte is the control field of the sent message and Response_code is the response code of the response message; check the global state transition graph, and if the same state identifier does not exist, create a new state node; where the global state transition tree is a directed graph gradually constructed by the system during fuzz testing, and initially only the initial state node 0 represents the initial state of the protocol entity program to be tested;
[0103] S2.2.2: Define and construct a global state transition graph. The system maintains a global directed graph structure to represent the state space of the protocol entity program to be tested; each state node v is uniquely identified by a tuple 〈Control_Byte,Response_code〉, where Control_Byte is the message control field that triggers the state, and Response_code is the corresponding response code; the state node also records the cumulative number of times it has been accessed, and maintains a seed set that can reach the state; and node 0 represents the initial state of the protocol entity program to be tested;
[0104] S2.2.3: Maintain the state transition relationship; perform timing processing on the message sequence in the test seed; for each group of sent messages and response messages, construct its state identification tuple 〈Control_Byte, Response_code〉; therefore, one test seed corresponds to several state identification tuples 〈Control_Byte, Response_code〉; check each state identification tuple 〈Control_Byte, Response_code〉, and if it has not been recorded, create a corresponding state node in the global state transition graph; then the system establishes directed edges between each pair of adjacent state nodes, and updates the statistical information such as the node's access count, seed set, and transition frequency.
[0105] For all message sequences in the seed, each group of sent messages and response messages is processed in time sequence to construct a corresponding two-tuple 〈Control_Byte, Response_code〉 state identifier; the system performs the following operations on each state identifier: first check and ensure that the state identifier has a corresponding node in the global state transition graph, and update the access count of the node; then for each pair of consecutive state identifiers, a directed edge is established between the previous state node and the next state node, and the frequency statistics of the transition to the next state node are recorded and updated in the previous state node.
[0106] Taking the RTSP protocol as an example, consider a test sequence containing four messages: first, send an OPTIONS message to establish a connection and get a response code of 200; then send a SETUP message to configure transmission parameters and get a response code of 200; then send a PLAY message to start playing and get a response code of 200; finally send the mutated message "SETUPabcde" and get an error response code of 405. The system converts this sequence into a state transition path: 0-><OPTIONS,200> -><SETUP,200> -><PLAY,200> -><UnKnow,405> In this process, the system not only creates the corresponding state nodes, but also establishes directed edges that reflect the state transition relationship, and records the number of visits to each node and the transition frequency of the edge;
[0107] S2.3: Determine the seed persistence strategy; persist or clear the current test seed according to its execution effect. For the test seed that needs to be saved, add it to the persistent seed library and associate it with the seed set of all nodes on its state transfer path.
[0108] S2.3.1: Determine whether the seed needs to be saved; check whether a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time during the seed execution process. When a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time, the seed is marked as a state to be saved;
[0109] S2.3.2: Execute seed persistence; for the marked seeds, add their complete contents to the persistent seed library for storage; at the same time, associate the seed with all state nodes on its state transfer path, ensuring that the seed set of each relevant node contains the seed. Through this association mechanism, a dedicated seed set is established and maintained for each state node. The seeds in the seed set have the ability to set the state of the protocol entity program under test to the corresponding state by replaying part of the message sequence.
[0110] like Figure 2 As shown, S3.1: Execute state replay and detection; for the state nodes that need to be subdivided, replay the seed set for testing; replay the message in the seed until the state of the protocol entity program to be tested is set to the current state, calculate the coverage edge information of the current execution, and then send a specific detection message to collect new execution coverage edge information, and then calculate the path increment feature and response code;
[0111] S3: Double-layer state segmentation stage; select the seeds saved in each state, form a corresponding state seed set and replay the relevant message sequence to guide the server to the target state, then send a specially constructed detection message and collect the response code and execution path feedback information generated in the test; preliminarily classify the seeds according to the response code, and then use the local sensitive hashing algorithm to analyze the path characteristics of each seed; by comparing the similarity of the path increments, the seeds with similar path characteristics are classified into the same refined state;
[0112] S3.1.1: Select a subdivision state node; periodically monitor the global state node, and when the number of seeds that have not been subdivided in a state node exceeds a preset threshold of 50 seeds, mark it as a state to be subdivided;
[0113] S3.1.2: Replay seeds: For each seed to be subdivided, replay its message sequence until the protocol entity program to be tested is guided to the current state; record the complete coverage edge information at this time as the reference path P_pre, and then send a message of the same special construction to all seeds, at this time obtain the current complete coverage edge information P_cur and the current response code information, and then use the reference path P_pre and the complete coverage edge information P_cur to calculate the incremental path P_m; the specially constructed message is taken from the random message of the random initial seed, and the message of the initial seed should be guaranteed to be legal as much as possible.
[0114] S3.1.3: Calculate the path increment; analyze the difference between the reference path P_pre and the complete coverage edge information P_cur, and calculate the incremental path P_m;
[0115] Create an empty bitmap array, and then compare each byte in the baseline path P_pre and the complete coverage edge information P_cur one by one; when the i-th byte is found to be different, the system sets the (i&7)th bit of the (i>>3)th byte in the bitmap to 1, and obtains the incremental path P_m; in this way, the system encodes the path difference into a compact binary bitmap, providing a basis for subsequent similarity analysis. The algorithm is formally described as:
[0116] P m [i>>3]|=(1<<(i&7)),if P cur [i]≠P pre [i]
[0117] Among them, i>>3 means shifting i right by 3 bits, and the operation is used to calculate the index of the target byte in the bitmap. i&7 means AND operation, and the operation is used to determine the position of the specific bit operation;
[0118] S3.2: Perform the first level division based on the response code; analyze the response codes of all seeds to the same special detection message, and classify the seeds with the same response code into the same sub-state, thereby completing the first level of state segmentation; Specifically, the system analyzes the response codes of all seeds to the same special detection message, and classifies the seeds with the same response code into the same sub-state, thereby completing the first level of state segmentation. This division mechanism is based on the following principle: when two seeds are identified as the same basic state but generate different response codes for the same detection message, it means that the two seeds actually correspond to different internal states of the protocol entity program to be tested. Therefore, the system realizes the preliminary segmentation of the state space through the difference in response codes.
[0119] S3.3: Perform the second-level partitioning based on the incremental path. For the state seed set in each sub-state, the second-level partitioning is performed using the local sensitive hashing algorithm according to the incremental path P_m calculated in S3.1.
[0120] The main advantage of introducing the local sensitive hashing algorithm is that the traditional path similarity analysis requires calculating the Hamming distance for each pair of seeds, and its computational complexity increases quadratically with the number of seeds, resulting in significant performance overhead in large-scale testing scenarios. The local sensitive hashing algorithm uses a voting mechanism of multiple hash buckets to map seeds of similar paths to the same bucket with a higher probability, allowing the system to complete similarity clustering in linear time. At the same time, by adjusting the number of hash buckets and the similarity threshold, the system can flexibly balance the accuracy and computational efficiency of clustering;
[0121] S4: State evaluation and selection stage: conduct a comprehensive evaluation of the state nodes, systematically analyze the topological structure of the global state transition graph to identify error state nodes and key control nodes, and conduct a comprehensive analysis based on the node's in-degree characteristics, transition probability and improved betweenness centrality index;
[0122] S4.1: Global state evaluation and selection; Analyze the characteristics of each node in the global state transition graph and evaluate and score it; Before each round of fuzz testing, comprehensively score the node based on its topological characteristics, historical benefits, and state characteristics, and select the appropriate state for the next round of testing;
[0123] S4.1.1: Simple state selection; when the number of fuzz tests is less than 10,000, a simple state selection strategy is used, that is, polling and selecting each node in the current global state graph to ensure that the state space of the entity program of the protocol to be tested is fully explored in the early stage of fuzz testing; otherwise, go to step S4.1.2;
[0124] S4.1.2: Calculate time-decay benefits; for each state node, introduce a time-decay mechanism to evaluate its historical test benefits; this mechanism gives historical discoveries a weight that decreases over time, so that recent test discoveries have a higher evaluation value. This design is based on the following considerations: Although early discoveries have opened up basic test space, their reference value for current decisions will decrease over time; in contrast, recent discoveries often reflect deeper program behaviors, and these new discoveries will become breakthroughs for further exploration. Through this dynamic weight adjustment mechanism, the system can more accurately evaluate the current value of state nodes, thereby optimizing the allocation of test resources. This mechanism gives historical discoveries a weight that decreases over time, so that recent test discoveries have a higher evaluation value. The specific calculation formula is:
[0125]
[0126] Among them, γ takes the value of 0.95 as the time decay coefficient, n is the current test round, t is the historical round index, paths_discovered(v,t) represents the number of new seeds discovered by state node v in the tth test round, and selected_times(v) represents the total number of times state node v is selected for fuzz testing;
[0127] S4.1.3: Identify error states; count the in-degree of each state node, calculate the number of transfer edges from all other state nodes to the node, and divide it by the total number of state nodes to get the in-degree ratio; when the in-degree ratio of a state node exceeds 0.7, the node is judged as a potential error state, and its score weight is reduced to 0.3 times, thereby reducing the tendency to test such states; this judgment is based on the following observations: In protocol implementations, error handling usually adopts a unified termination logic, so error states often become the convergence point of a large number of state transfers. Although this type of state has a high in-degree, its internal logic is relatively simple, and it is difficult to discover new program behaviors by continuing testing. Therefore, the system reduces its score weight to 0.3 times and allocates more testing resources to normal state nodes that may contain complex interaction logic.
[0128] S4.1.4: Analyze the transition probability; the system calculates the transition frequency between all state nodes, and assigns scoring rewards to nodes based on the rarity of state node transitions. This design is based on the following principles: In mutation-based fuzz testing, due to the randomness of the mutation operation and the complexity of the message structure, the probability of a mutated message generating an erroneous response is much higher than the probability of generating a legitimate message. Therefore, those state paths that still maintain a low transition frequency during the test process can better reflect the correct state transition behavior. By improving the scores of the state nodes corresponding to these rare paths, the system can better guide test resources to invest in these valuable exploration directions. Calculate the transition frequency between all state nodes, count all possible transfer targets and the corresponding number of transfers for the state node v, and then calculate its transition probability score; the specific calculation formula is:
[0129]
[0130] Among them, V is the set of state nodes, trans(i,v) represents the number of historical transfers from state node i to state node v, Represents the total number of transitions to all sub-state nodes to which state node i can transfer;
[0131] S4.1.5: Evaluate node betweenness centrality; introduce betweenness centrality calculation to evaluate the importance of a state node in the global state transition graph. Update the betweenness centrality score of the node every 10 rounds of state selection. Quantify its importance by calculating the frequency of the state node appearing in the shortest path from the initial state to other states. The specific calculation formula is:
[0132]
[0133] Among them, σst(v) represents the number of paths containing node v in the shortest path from the initial state 0 to any other state, and σst represents the total number of shortest paths from the initial state to other states, that is, the total number of state nodes;
[0134] S4.1.6: Final score calculation: Taking into account the evaluation results of the above dimensions, the final score of each state node is calculated. The specific calculation formula is:
[0135] Score(v)=R(v)·α(v)·(1+BC(v))+Score trans (v)
[0136] Where R(v) is the time-decayed historical return value calculated in the previous step. The coefficient α(v) takes the value of 0.3 only when the in-degree ratio of a node exceeds 0.7, otherwise it takes the value of 1.0. BC(v) represents the improved betweenness centrality score of the node, which is used to reflect the importance of the node in state transition. Score_trans(v) is the transition probability score of the node.
[0137] After obtaining the original scores of all nodes, normalization is performed: the sum of all node scores is calculated, and then the score of each node is divided by the sum to obtain a normalized probability distribution; based on this probability distribution, a roulette wheel selection algorithm is used to determine the target state for the next round of testing, so that states with higher scores have a greater probability of being selected, while ensuring that states with lower scores still have a chance to be explored;
[0138] S4.2: Specific subdivision state seed selection; check whether there are new seeds that have not been fuzz tested in the selected state, if so, give priority to these new seeds for testing; if there are no new seeds, the system randomly selects a subdivision state in the state based on uniform probability distribution, and then randomly selects a seed in the selected subdivision state as the test object; after completing the seed selection, determine the specific message for the fuzz test; analyze the complete state transition path of the seed, and determine the message sequence required to guide the protocol entity program to be tested to the target state, and then select the next message corresponding to the state transition point in the sequence as the fuzz test object; pass the selected seed and the specific message location information that needs to be fuzz tested to step S1 as the input parameter for a new round of fuzz testing.
[0139] S5: Fuzz test execution phase; select the target state and seed according to the evaluation result of step S4, and perform random mutation on them to continuously generate new seeds; use the mutated new seeds to execute the entity program of the protocol to be tested, and collect the coverage information and response message sequence during its execution process, and then use the analysis mechanism established in steps S2 and S3, including state identification extraction and path feature analysis, to characterize the state characteristics of the new execution results;
[0140] S5.1: Enter the formal fuzz testing phase; in each round of testing, select the target state and its corresponding seed for testing based on the evaluation results of step S4;
[0141] S5.2: Generate new seeds for the current test seeds using a mutation strategy based on a genetic algorithm;
[0142] S5.2.1: Determine the fuzzy area; since a single seed is composed of multiple messages, this stage first obtains the target seed to be tested and the specific message position that needs to be fuzz tested through the interface of step S4; in the subsequent mutation process, the specific message is mutated and the mutated message is replaced back to the original position, thereby constructing a new test seed;
[0143] S5.2.2: Enter the deterministic mutation stage; perform mutation operations on each byte position of the target message in sequence according to a preset fixed order, including the steps of bit-by-bit flipping, byte-by-byte flipping, and byte-by-byte replacement of special values;
[0144] S5.2.3: Enter the non-deterministic mutation stage; randomly select and combine multiple mutation operators to mutate the target message. The mutation operators include random position insertion, random block deletion, and block splicing and reorganization. In each round of mutation, the operator combination and mutation position used are randomly determined to generate a new test message.
[0145] S5.3: Start the protocol entity program to be tested, send the mutated new seed for testing, collect the execution results, including program coverage data and the complete response message returned by the protocol entity program to be tested, and pass the feedback information to step S2 through the interface for analysis;
[0146] S5.3.1: Start the protocol entity program to be tested. The seed to be sent consists of the mutated target message and other original messages that have not been mutated in sequence. Establish a network connection with the protocol entity program to be tested and send the messages one by one according to the timing of the messages. For each message, after the message is sent, periodically check the edge coverage feedback of the protocol entity program to be tested. When the edge coverage information of two consecutive checks is consistent, it is determined that the message has been completely processed, and then the next message in the sequence is sent until all the messages in the seed are sent.
[0147] S5.3.2: When all messages in the seed have been sent and responses received, the system organizes the test execution results, including the final coverage edge information and the complete response message sequence, and passes this feedback information to step S2 through the interface for analysis;
[0148] S5.4: Repeat steps S5.2 to S5.3 until the preset upper limit of mutation times is reached; then return to step S5.1 to select a new state and seed for the next round of fuzzy testing.
[0149] S6: Feedback analysis phase; systematically analyze the test execution results and maintain global status information.
[0150] For this embodiment (referred to as SstateFuzzer in the table), in order to prove the effectiveness of the method of the present invention, the present invention sets up a 24-hour fuzz test in the same experimental environment. The experiment selects the most advanced stateful fuzz test tool AFL-Net as the control group, and the test set selects the widely used open source protocols RTSP and DTLS. The same protocol entity program to be tested is tested, and each group of experiments lasts for 24 hours. In order to reduce the impact of randomness in the fuzz test, all experiments are repeated three times, and the results are averaged. The experimental results are shown in the following table:
[0151]
[0152] As can be seen from the above table, on the two widely used protocols RTSP and DTLS, the code coverage of the present invention is improved by 2.4% and 6.7% respectively compared with the benchmark tool AFL-NET, and the number of vulnerabilities discovered is increased by 16% and 41% respectively, which fully demonstrates the significant advantages of the mechanism in state identification and test guidance.
[0153] Any of the methods or steps described above may be stored as computer instructions or programs in various types of computer memories, and the computer instructions or programs may be recognized by various types of computer processors to implement any of the methods or steps described above.
[0154] The embodiments described above are only descriptions of the preferred modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A protocol fuzz testing method based on fine-grained state partitioning and selection, characterized by: The method comprises the following steps: S1: fuzz testing preprocessing stage; Deploy the entity program of the protocol to be tested and configure the initial seed set; S2: Basic state identification stage; receive complete feedback information of each initial seed in the initial seed set, including the final coverage edge information and the response message sequence returned by the protocol entity program to be tested, establish the correspondence between the sent message and the response message, then extract the response code from the response message, and construct a joint state identification in combination with the control field of the corresponding sent message, and use the information to build and maintain the global state transition graph; and use all the initial seeds in the initial seed set as test seeds; S3: Double-layer state segmentation stage; select the seeds saved in each state, form a corresponding state seed set and replay the relevant message sequence to guide the server to the target state, then send a specially constructed detection message and collect the response code and execution path feedback information generated in the test; preliminarily classify the seeds according to the response code, and then use the local sensitive hashing algorithm to analyze the path characteristics of each seed; by comparing the similarity of the path increments, the seeds with similar path characteristics are classified into the same refined state; S4: state assessment and selection stage; Comprehensively evaluate the state nodes, systematically analyze the topological structure of the global state transition graph to identify error state nodes and key control nodes, and conduct a comprehensive analysis based on the node's in-degree characteristics, transition probability, and improved betweenness centrality index; S5: Fuzz test execution phase; select the target state and seed according to the evaluation result of step S4, and perform random mutation on them to continuously generate new seeds; use the mutated new seeds to execute the entity program of the protocol to be tested, and collect the coverage information and response message sequence during its execution process, and then use the analysis mechanism established in steps S2 and S3, including state identification extraction and path feature analysis, to characterize the state characteristics of the new execution results; S6: Feedback analysis phase; systematically analyze the test execution results and maintain global status information.
2. A protocol fuzz testing method based on fine-grained state division and selection according to claim 1, characterized in that: The specific steps of step S1 include: S1.1: Deploy the complete file of the protocol entity program to be tested, configure the environment and dependencies required by the protocol entity program to be tested, and compile it using the afl-gcc stub compilation tool; S1.1.1: Correctly deploy the complete file of the protocol entity program to be tested in the local environment computer, and configure the environment and dependencies required for its operation; use the compilation tool to pre-compile the protocol entity program to be tested, and specifically identify and distinguish each basic block in the code during the compilation process; a basic block refers to a linearly executed instruction sequence in the program, which has the characteristics of a single entry point and a single exit point; S1.1.2: Use the instrumentation compilation tool to instrument the protocol entity program to be tested, insert monitoring code at the entrance of each basic code block, and assign a unique random identification value to each basic block; when the program execution transfers from one basic block to another, the random identification values of the two basic blocks before and after are combined into an edge identifier; by recording and counting these edge identifiers, the code paths covered during program execution can be accurately tracked, providing the necessary coverage feedback for subsequent analysis, where "edge" and "covered edge" specifically refer to the jump behavior of the basic blocks monitored when the protocol entity program to be tested is running; S1.2: Prepare the initial seed library; locally start the executable file of the entity program of the protocol to be tested, and perform multiple rounds of normal communication with it by simulating the client locally; during the communication process, use the network packet analysis tool to capture the communication traffic and extract the message data sent by the client as the initial seed; start the fuzz test with at least one legal initial seed; S1.2.1: Start a network packet analysis tool locally and continuously monitor the local network interface to capture communication traffic; S1.2.2: Start a protocol entity program to be tested locally as a server, and start a client program separately; simulate the normal communication process for testing according to the RFC specification document; during this process, the network packet analysis tool captures the message sequence sent by the client, combines the captured multiple messages in chronological order and splits them with specific delimiters to form a usable initial seed; S1.2.3: As needed, repeat step S1.2.2 to obtain initial seeds in multiple different scenarios to form an initial seed set.
3. A protocol fuzz testing method based on fine-grained state division and selection according to claim 2, characterized in that: The specific steps of step S2 include: S2.1: Receive the complete feedback information of each initial seed in step S1, including the final coverage edge information and the response message sequence returned by the protocol entity program under test; identify the response code of each response message and extract the control field information from the corresponding sent message; S2.2: Construct a state identifier and update the state transition information; construct a state identifier based on the extracted response code and control field information, and use them to construct and maintain relevant information of the global state transition graph; S2.3: Determine the seed persistence strategy; persist or clear the current test seed according to its execution effect. For the test seed that needs to be saved, add it to the persistent seed library and associate it with the seed set of all nodes on its state transfer path.
4. A protocol fuzz testing method based on fine-grained state division and selection according to claim 3, characterized in that: The specific steps of step S2.1 include: S2.1.1: Establish a correspondence between sent messages and response messages; each sent message corresponds to a response message; S2.1.2: Get the response code of each response message. Execute the callback function GetCode for each response message. The callback function GetCode extracts and processes certain fields from any response message, and finally hashes them into an integer variable, and regards the integer variable as the response code of the response message; S2.1.3: Take out the control field of each sent message; execute the callback function GetControl for each message in the test seed. The callback function GetControl extracts certain fields from any sent message and processes them, and finally returns a character variable, and regards the character variable as the control field of the sent message.
5. A protocol fuzz testing method based on fine-grained state division and selection according to claim 4, characterized in that: The specific steps of step S2.2 include: S2.2.1: Construct a state joint identifier; for each set of sent messages and corresponding response messages, construct a two-tuple 〈Control_Byte,Response_code〉 as a state identifier, where Control_Byte is the control field of the sent message, and Response_code is the response code of the response message; S2.2.2: Define and construct a global state transition graph. The system maintains a global directed graph structure to represent the state space of the protocol entity program to be tested; each state node v is uniquely identified by a tuple 〈Control_Byte,Response_code〉, where Control_Byte is the message control field that triggers the state, and Response_code is the corresponding response code; the state node also records the cumulative number of times it has been accessed, and maintains a seed set that can reach the state; and node 0 represents the initial state of the protocol entity program to be tested; S2.2.3: Maintain the state transition relationship; perform timing processing on the message sequence in the test seed; for each group of sent messages and response messages, construct its state identification tuple 〈Control_Byte, Response_code〉; therefore, one test seed corresponds to several state identification tuples 〈Control_Byte, Response_code〉; check each state identification tuple 〈Control_Byte, Response_code〉, and if it has not been recorded, create a corresponding state node in the global state transition graph; then the system establishes directed edges between each pair of adjacent state nodes, and updates the statistical information such as the node's access count, seed set, and transition frequency.
6. A protocol fuzz testing method based on fine-grained state division and selection according to claim 5, characterized in that: The specific steps of step S2.3 include: S2.3.1: Determine whether the test seed needs to be saved; check whether a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time during the execution of the test seed. When a new coverage edge is triggered for the first time, a new state node is discovered for the first time, or a new state transition path is generated for the first time, the test seed is marked as a state to be saved; S2.3.2: Execute seed persistence; for the marked test seeds, add their complete contents to the persistent seed library for storage; at the same time, associate the test seed with all state nodes on its state transfer path, ensuring that the seed set of each relevant node contains the test seed. Through this association mechanism, a dedicated seed set is established and maintained for each state node. The seeds in the seed set have the ability to set the state of the entity program of the protocol to be tested to the corresponding state by replaying part of the message sequence.
7. A protocol fuzz testing method based on fine-grained state division and selection according to claim 6, characterized in that: The specific steps of step S3.1 include: S3.1: Execute state replay and detection; for the state nodes that need to be subdivided, replay the seed set for testing; replay the messages in the seed until the state of the protocol entity program to be tested is set to the current state, calculate the coverage edge information of the current execution, and then send a specific detection message to collect new execution coverage edge information, and then calculate the path increment feature and response code; S3.1.1: Select a subdivision state node; periodically monitor the global state node, and when the number of seeds that have not been subdivided in a state node exceeds a preset threshold of 50 seeds, mark it as a state to be subdivided; S3.1.2: Replay seeds: For each seed in the state to be subdivided, replay its message sequence until the protocol entity program to be tested is guided to the current state; record the complete coverage edge information at this time as the reference path P_pre, and then send a message of the same special structure to all seeds, at this time obtain the current complete coverage edge information P_cur and the current response code information, and then use the reference path P_pre and the complete coverage edge information P_cur to calculate the incremental path P_m; S3.1.3: Calculate the path increment; analyze the difference between the reference path P_pre and the complete coverage edge information P_cur, and calculate the incremental path P_m; Create an empty bitmap array, and then compare each byte in the reference path P_pre and the complete coverage edge information P_cur one by one; when the i-th byte is found to be different, the system sets the (i&7)th bit of the (i>>3)th byte in the bitmap to 1, and obtains the incremental path P_m; S3.2: Perform the first level division based on the response code; analyze the response codes of all seeds to the same special detection message, and classify the seeds with the same response code into the same sub-state, thereby completing the first level of state subdivision; S3.3: Perform the second-level partitioning based on the incremental path. For the state seed set in each sub-state, the second-level partitioning is performed using the local sensitive hashing algorithm according to the incremental path P_m calculated in S3.
1.
8. A protocol fuzz testing method based on fine-grained state division and selection according to claim 7, characterized in that: The specific steps of step S4 include: S4.1: Global state evaluation and selection; Analyze the characteristics of each node in the global state transition graph and evaluate and score it; Before each round of fuzz testing, comprehensively score the node based on its topological characteristics, historical benefits, and state characteristics, and select the appropriate state for the next round of testing; S4.1.1: Simple state selection; when the number of fuzz tests is less than 10,000, a simple state selection strategy is used, that is, polling and selecting each node in the current global state graph to ensure that the state space of the entity program of the protocol to be tested is fully explored in the early stage of fuzz testing; otherwise, go to step S4.1.2; S4.1.2: Calculate time decay benefits. For each state node, introduce a time decay mechanism to evaluate its historical test benefits. This mechanism gives recent test findings a higher evaluation value by giving historical findings a weight that decreases over time. The specific calculation formula is: Among them, γ takes the value of 0.95 as the time decay coefficient, n is the current test round, t is the historical round index, paths_discovered(v,t) represents the number of new seeds discovered by state node v in the tth test round, and selected_times(v) represents the total number of times state node v is selected for fuzz testing; S4.1.3: Identify error states; count the in-degree of each state node, calculate the number of transition edges from all other state nodes to the node, and divide it by the total number of state nodes to get the in-degree ratio; when the in-degree ratio of a state node exceeds 0.7, the node is judged as a potential error state, and its score weight is reduced to 0.3 times, thereby reducing the tendency to test such states; S4.1.4: Analyze the transition probability; calculate the transition frequency between all state nodes, count all possible transition targets of state node v and the corresponding number of transitions, and then calculate its transition probability score; the specific calculation formula is: Among them, V is the set of state nodes, trans(i,v) represents the number of historical transfers from state node i to state node v, Represents the total number of transitions to all sub-state nodes to which state node i can transfer; S4.1.5: Evaluate node betweenness centrality; introduce betweenness centrality calculation to evaluate the importance of a state node in the global state transition graph. Update the betweenness centrality score of the node every 10 rounds of state selection. Quantify its importance by calculating the frequency of the state node appearing in the shortest path from the initial state to other states. The specific calculation formula is: Among them, σst(v) represents the number of paths containing node v in the shortest path from the initial state 0 to any other state, and σst represents the total number of shortest paths from the initial state to other states, that is, the total number of state nodes; S4.1.6: Final score calculation: Taking into account the evaluation results of the above dimensions, the final score of each state node is calculated. The specific calculation formula is: Score(v)=R(v)·α(v)·(1+BC(v))+Score trans (v) Where R(v) is the time-decayed historical return value calculated in the previous step. The coefficient α(v) takes the value of 0.3 only when the in-degree ratio of a node exceeds 0.7, otherwise it takes the value of 1.
0. BC(v) represents the improved betweenness centrality score of the node, which is used to reflect the importance of the node in state transition. Score_trans(v) is the transition probability score of the node. After obtaining the original scores of all nodes, normalization is performed: the sum of all node scores is calculated, and then the score of each node is divided by the sum to obtain a normalized probability distribution; based on this probability distribution, a roulette wheel selection algorithm is used to determine the target state for the next round of testing, so that states with higher scores have a greater probability of being selected, while ensuring that states with lower scores still have a chance to be explored; S4.2: Specific subdivision state seed selection; check whether there are new seeds that have not been fuzz tested in the selected state, if so, give priority to these new seeds for testing; if there are no new seeds, the system randomly selects a subdivision state in the state based on uniform probability distribution, and then randomly selects a seed in the selected subdivision state as the test object; after completing the seed selection, determine the specific message for the fuzz test; analyze the complete state transition path of the seed, and determine the message sequence required to guide the protocol entity program to be tested to the target state, and then select the next message corresponding to the state transition point in the sequence as the fuzz test object; pass the selected seed and the specific message location information that needs to be fuzz tested to step S1 as the input parameter for a new round of fuzz testing.
9. A protocol fuzz testing method based on fine-grained state division and selection according to claim 8, characterized in that: The specific steps of step S5 include: S5.1: Enter the formal fuzz testing phase; in each round of testing, select the target state and its corresponding seed for testing based on the evaluation results of step S4; S5.2: Generate test seeds for the currently selected seeds using a mutation strategy based on a genetic algorithm; S5.2.1: Determine the fuzzy area; since a single seed is composed of multiple messages, this stage first obtains the target seed to be tested and the specific message position that needs to be fuzz tested through the interface of step S4; in the subsequent mutation process, the specific message is mutated and the mutated message is replaced back to the original position, thereby constructing a new test seed; S5.2.2: Enter the deterministic mutation stage; perform mutation operations on each byte position of the target message in sequence according to a preset fixed order, including the steps of bit-by-bit flipping, byte-by-byte flipping, and byte-by-byte replacement of special values; S5.2.3: Enter the non-deterministic mutation stage; randomly select and combine multiple mutation operators to mutate the target message. The mutation operators include random position insertion, random block deletion, and block splicing and reorganization. In each round of mutation, the operator combination and mutation position used are randomly determined to generate a new test message. S5.3: Start the protocol entity program to be tested, send the mutated test seed for testing, collect the execution results, including program coverage data and the complete response message returned by the protocol entity program to be tested, and pass the feedback information to step S2 through the interface for analysis; S5.3.1: Start the protocol entity program to be tested. The test seed to be sent consists of the mutated target message and other original messages that have not been mutated in sequence. Establish a network connection with the protocol entity program to be tested and send the messages one by one according to the timing of the messages. For each message, after the message is sent, periodically check the edge coverage feedback of the protocol entity program to be tested. When the edge coverage information of two consecutive checks is consistent, it is determined that the message has been completely processed, and then the next message in the sequence is sent until all the messages in the test seed are sent. S5.3.2: When all messages in the test seed have been sent and responses received, the system organizes the test execution results, including the final coverage edge information and the complete response message sequence, and passes this feedback information to step S2 through the interface for analysis; S5.4: Repeat steps S5.2 to S5.3 until the preset upper limit of mutation times is reached; then return to step S5.1 to select a new state and seed for the next round of fuzzy testing.
10. A system comprising the protocol fuzz testing method based on fine-grained state partitioning and selection according to claim 1, characterized in that: It includes a state division module, a state evaluation module, a test execution module and a feedback analysis module; each module works together to achieve efficient fuzz testing of the entity program of the protocol to be tested; Fuzz test preprocessing module; responsible for deploying the entity program of the protocol to be tested; State segmentation module: This module is responsible for implementing a two-layer state recognition mechanism, including a basic feature extraction unit and a state segmentation processing unit. The basic feature extraction unit is responsible for extracting the response code and control field information from the interaction process and constructing the initial state identifier. The state segmentation processing unit works regularly to further segment the state through path analysis and similarity clustering. State evaluation module: This module is responsible for systematically analyzing and evaluating the global state transition graph and guiding the allocation of fuzz testing resources. In the early stage of fuzz testing, the module selects state nodes by random polling to ensure full exploration of the initial state space. When enough test samples are accumulated, the module switches to the heuristic evaluation mechanism. By analyzing the topological structure of the state transition graph, the module identifies the error state nodes and key control nodes, and conducts a comprehensive analysis based on the node's in-degree characteristics, transition probability and improved betweenness centrality index. In the scoring process, the module introduces a time decay mechanism to calculate the node's historical benefits, and combines it with the node feature score, and finally generates a score for each state node that reflects its importance. Through this mechanism that combines random exploration and heuristic evaluation, the module provides a reliable decision-making basis for the reasonable allocation of test resources. Test execution module: This module is responsible for executing the specific fuzz testing process. Based on the target state and seed provided by the state evaluation module, this module performs mutation testing on the target message specified therein. During the mutation process, a new test seed is generated by randomly combining multiple mutation operators while keeping other messages in the seed unchanged. For each mutated new seed, the module starts the entity program of the protocol to be tested and sends a sequence of messages in order. Feedback analysis module: This module is responsible for systematically analyzing the test execution results and maintaining the global state information; this module collects the program execution path and response characteristics of each test, analyzes the state transition sequence and updates the statistical data of node access frequency and transition probability in the global state transition graph.
Citation Information
Patent Citations
Gray box text protocol fuzz testing method and system based on field perception
CN117478566A
Network protocol fuzz testing method based on state transition relation and mutation strategy optimization
CN118784340A
Protocol vulnerability mining test method and system based on fine-grained state guidance
CN116962262A
Network protocol fuzz testing method based on refined mutation probability learning
CN119363635A
Cited By
Real-time threat monitoring and defending method and system for digital infrastructure
CN120200849A
Network protocol fuzzy test seed evaluation method based on multi-standard combination weighting
CN121603404A