Protocol fuzz testing method based on program variation

By using instrumentation and scheduler management at compile time of protocol client, encrypted and effective test cases are generated, which solves the problem of low protocol state coverage in the existing technology, and achieves more in-depth protocol testing and vulnerability discovery.

CN120371703APending Publication Date: 2025-07-25SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471958.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing fuzz testing technology is difficult to generate encrypted and effective test cases in protocol testing, and cannot effectively test the decrypted logic, resulting in low protocol state coverage and branch coverage, and it is difficult to discover complex protocol control flow vulnerabilities.

Method used

The scope of variation is determined through static analysis, and the data flow and control flow instructions of the protocol client are inserted at compile time to generate encrypted and effective test cases, and the communication between the variant client and the system under test is managed through the scheduler, and coverage feedback is used to guide deeper protocol state exploration.

Benefits of technology

The protocol state coverage and branch coverage were significantly improved, and multiple unknown vulnerabilities were discovered, with high testing efficiency and versatility, and able to effectively test decrypted logic and complex protocol scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371703A_ABST
    Figure CN120371703A_ABST
Patent Text Reader

Abstract

The invention relates to a protocol fuzz testing method based on instruction level variation. The protocol fuzz testing method mainly solves the problem of protocol state accessibility of an existing fuzz testing technology in protocol testing. Aiming at the characteristics of a protocol, the invention provides a three-stage fuzzy test framework, which comprises the following steps of: firstly, determining a variation range through static analysis (1), then carrying out instrumentation on data flow and control flow instruction operations in a program during compiling, and finally, managing communication between a variation client and a tested system through a scheduler (2); and (4) guiding deeper protocol state exploration by utilizing coverage rate feedback. According to the method, the traditional fuzzy test technology is expanded, the problem that the encryption structure is damaged by the traditional message-level variation in the protocol is solved by introducing the instruction-level variation, and the protocol state coverage rate and the branch coverage rate are remarkably improved. Experimental results show that the method discovers a plurality of unknown vulnerabilities in the implementation of a plurality of protocols, and has relatively high test efficiency and universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a protocol fuzz testing technology, belonging to the technical fields of network security and software testing. A protocol fuzz testing method based on instruction-level mutation is proposed. By means of static analysis, mutation point instrumentation, and scheduling fuzz testing, the state reachability problem in existing fuzz testing technologies for protocol testing is solved, and the decrypted logic of the protocol can be effectively tested to discover more potential vulnerabilities. Background Art

[0002] Protocols (such as TLS) are the cornerstone of network communication security and are widely used to protect the confidentiality, integrity, and authenticity of data transmission. However, the design and implementation of protocols are complex and prone to introducing vulnerabilities. Once exploited by attackers, it may lead to serious security consequences. For example, the "Heartbleed" vulnerability in OpenSSL allows attackers to read sensitive information in the server memory, affecting more than 17% of servers globally. Therefore, identifying and fixing vulnerabilities in protocols is crucial for ensuring network security. Fuzz testing is an efficient software testing technology that discovers potential vulnerabilities by generating a large number of random or semi-random test cases to trigger abnormal behaviors of the system under test. Traditional fuzz testing technologies face challenges in protocol testing. The messages in protocols are usually encrypted before transmission. Traditional message-level mutation methods will damage the encryption structure, resulting in test cases failing to pass the decryption stage, and thus the decrypted logic cannot be tested. For example, in the TLS1.3 protocol, most handshake messages are encrypted, and existing fuzz testing tools show low code coverage in TLS1.3 testing and cannot effectively discover vulnerabilities in the encrypted state.

[0003] Existing protocol fuzz testing methods are mainly divided into two categories: one is traditional fuzz testing tools based on message-level mutation, such as AFLNet and SGFuzz, which generate test cases by mutating message contents; the other is fuzz testing tools based on program code-level mutation, such as Fuzztruction and IoTFuzzer, which generate semi-effective test cases by mutating operations inside the program and can bypass encryption verification, but mainly target non-interactive applications or specific platforms (such as Android) and cannot fully support protocols for two-way communication.

[0004] Therefore, the technical problem to be solved by the present invention is how to generate test cases with valid encryption but mutated data through instruction-level mutation in protocol testing, ensure that the test cases can pass the decryption stage, and thus effectively test the decrypted logic of the protocol to discover more potential vulnerabilities. Summary of the Invention

[0005] The present invention proposes a protocol fuzz testing method based on instruction-level mutation. By static analysis, the mutation range is determined. During compilation, data flow and control flow instructions in the program are instrumented to generate test cases with valid encryption but mutated data. A scheduler is used to manage the communication between the mutation client and the system under test, and coverage feedback is utilized to guide deeper exploration of protocol states. Experimental results show that this method significantly improves protocol state coverage and branch coverage in multiple protocol implementations, discovers multiple unknown vulnerabilities, and has high test efficiency and generality.

[0006] To solve the above technical problems, the present invention provides a protocol fuzz testing method based on instruction-level mutation. The solution is divided into three steps: static analysis, mutation point instrumentation, and scheduling fuzz testing. The specific steps are as follows:

[0007] 1. Static analysis stage

[0008] The mutation range in the protocol client is determined through static analysis. Functions related to message construction and transmission are identified, and the session structure is extracted as the core basis for the mutation range. The specific sub-steps include:

[0009] (1) Message sending function identification: Analyze the client source code to identify functions that call system calls (such as write, send, sendto) as message sending functions.

[0010] (2) Session structure extraction: Extract the first parameter of the message sending function as the session structure, which contains key information required for protocol connection (such as encryption parameters, protocol version, session key, etc.).

[0011] (3) Session function screening: Automatically screen out functions with the session structure as a parameter through a static analysis tool (Clang / LLVM) as session functions to ensure that the mutation range covers the core logic of message construction, conversion, and transmission.

[0012] 2. Mutation point instrumentation stage

[0013] During compilation, data flow and control flow instructions in the client program are instrumented to generate mutation points, ensuring that the mutation occurs before encryption and generating test cases with valid encryption but mutated data. The specific sub-steps include:

[0014] (1) Data flow instruction instrumentation:

[0015] (a) Insert mutation logic after the load instruction and before the store instruction to ensure that the mutation takes effect before the data enters memory.

[0016] (b) Avoid mutating instructions with pointer type operands to prevent program crashes.

[0017] (c) Generate test cases that can trigger integer overflow or out-of-bounds access by using mutation strategies such as zero values, boundary values, and random values.

[0018] (2) Control flow instruction instrumentation:

[0019] (a) Mutate the compare (cmp) instruction to change the program execution path by reversing the comparison result;

[0020] (b) Mutate the branch (switch) instruction to generate an abnormal protocol sequence (such as skipping key authentication steps) by changing the branch condition, triggering control flow-related vulnerabilities;

[0021] 3. Scheduling fuzz testing phase

[0022] Manage the communication between the mutation client and the system under test (SUT) through a scheduler, use coverage feedback to guide the selection of mutation points, explore deeper protocol states, and discover potential vulnerabilities. Specifically, it includes the following sub-steps:

[0023] (1) Mutation point activation control:

[0024] (a) Assign a unique identifier to each mutation point and control its activation status through a global variable;

[0025] (b) Support activating multiple mutation points simultaneously to simulate complex protocol interaction scenarios.

[0026] (2) Coverage feedback drive:

[0027] (a) The scheduler first activates all mutation points, records the mutation points that can trigger new coverage, and uses them as the preferred mutation points (PMPs);

[0028] (b) In subsequent fuzz testing, preferentially select PMPs for activation and dynamically adjust the mutation combination according to coverage to cover more protocol states;

[0029] (c) Introduce a state awareness mechanism to accurately identify the current protocol state by tracking the memory snapshot of the protocol server or enumerating variable values, guiding the mutation direction.

[0030] (3) Bidirectional communication management:

[0031] (a) Design an interaction protocol to ensure that the bidirectional communication between the mutation client and the system under test can be fully executed;

[0032] (b) Reset the protocol state after each interaction to avoid the residual state affecting the independence of subsequent test cases.

[0033] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the protocol fuzz testing method based on instruction-level mutation is implemented.

[0034] A computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the protocol fuzz testing method based on instruction-level mutation is implemented.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] (1) Effectively solve the problem of state reachability in protocol testing

[0037] Existing fuzz testing tools mainly rely on message-level mutation, that is, directly modifying the data of communication messages. However, in modern cryptographic protocols such as TLS1.3, the key messages in the handshake process are all in an encrypted state, and message-level mutation usually destroys the ciphertext structure, resulting in decryption failure, and thus the core logic after protocol decryption (such as certificate verification, session key generation) cannot be effectively tested.

[0038] The present invention innovatively proposes an instruction-level mutation technology, which performs fine-grained mutation of data flow and control flow instructions inside the client program and before the message is encrypted, so as to generate protocol messages that can be correctly encrypted and decrypted but have abnormal content, significantly improving the problem of protocol state reachability, and enabling fuzz testing to effectively cover the logic after decryption for the first time.

[0039] (2) Significantly improve the coverage rate of protocol fuzz testing and expand the scope of vulnerability mining

[0040] Existing protocol fuzz testing tools (such as SGFuzz, Fuzztruction) have obvious deficiencies in state coverage and code branch coverage, especially in the encrypted message stage of cryptographic protocols.

[0041] Through the proposed instruction-level fault injection and scheduling mechanism, the present invention has achieved a 44.69% increase in state coverage rate and a 16.36% increase in branch coverage rate in the fuzz testing of the TLS1.3 protocol. Specifically, this method can effectively reach the encrypted message states that cannot be accessed by traditional tools (such as TLS_ST_SW_FINISHED in the OpenSSL implementation), realizing a more comprehensive and in-depth test of protocol logic, thereby greatly expanding the coverage range and depth of vulnerability mining.

[0042] (3) Support the simulation of complex protocol attack scenarios and improve the ability to discover vulnerabilities

[0043] The prior art usually simply focuses on data mutation, making it difficult to discover complex protocol control flow vulnerabilities (such as authentication bypass and abnormal state jump vulnerabilities), and it is not easy to implement dynamic two-way interaction management between the client and the server.

[0044] By designing a dynamic management mechanism for the scheduler, the present invention can automatically and dynamically combine multiple fault points, realize the combined mutation of control flow and data flow, and accurately simulate real malicious client attack scenarios.

[0045] (4) Static analysis is accurate, efficient, and the fault injection scope is clear

[0046] Traditional instruction-level mutation techniques (such as Fuzztruction) have problems of overly large fault injection scope and overly broad instrumentation. A large number of instrumentations reduce the efficiency of fuzz testing and may introduce a large number of invalid fault points.

[0047] The fault point pruning method proposed by the present invention performs static analysis based on LLVM IR, accurately identifies session structures and key session functions in protocol implementations, and only performs fault injection instrumentation within the function scope that is truly related to protocol message construction and transmission, ensuring that the instrumentation scope is both accurate and efficient. This innovative technology not only improves the execution efficiency of fuzz testing but also avoids the risk of program crashes that may be caused by meaningless instrumentation.

[0048] (5) It has good generality and applicability

[0049] Existing protocol fuzz testing tools usually focus on specific types of protocols or scenarios and have poor applicability to other protocols.

[0050] Through general static analysis and instrumentation methods, the present invention can automatically identify key protocol data structures and functions without knowledge of the specific protocol format. It has been successfully verified on multiple mainstream cryptographic protocols (OpenSSL, wolfSSL, MbedTLS, LibreSSL, MatrixSSL, GnuTLS) and non-cryptographic protocols (libcoap), demonstrating that this method has excellent generality and can be extended to more network protocols and application scenarios.

[0051] (6) It has been fully verified through actual experiments and has strong practical value

[0052] The technical solution of the present invention has been verified through actual experiments. 10 real vulnerabilities have been successfully discovered in the implementations of 6 mainstream cryptographic protocols, among which 4 are 0day vulnerabilities first discovered in the industry (such as CVE-2024-0901 of wolfSSL, CVE-2024-23744 of MbedTLS, etc.). This actual experimental result significantly proves the effectiveness, practicality, and technical innovation of this technical solution, and has obvious technical advantages and practical value compared with the prior art. Brief Description of the Drawings

[0053] Figure 1 This is a schematic diagram for the mode test of this application;

[0054] Figure 2 This is an architecture diagram of the fuzz testing framework of this application;

[0055] Figure 3 This is the algorithm for the static analysis stage of this application;

[0056] Figure 4 This is the scheduling algorithm for the fuzz testing stage of this application;

[0057] Figure 5 This is the comparison of the coverage rate between this application and other methods in fuzz testing. Detailed Implementation Modes

[0058] The following elaborates in detail on the specific implementation modes of the present invention in conjunction with the attached drawings, but the protection scope of the present invention is not limited to the described embodiments.

[0059] Embodiment: The present invention provides a protocol fuzz testing method based on instruction-level mutation. This method inserts mutation points into the protocol client program and uses coverage feedback to guide fuzz testing to achieve in-depth testing of the logic after protocol decryption to discover potential vulnerabilities in the protocol implementation. This method includes five main stages: static analysis, mutation point instrumentation, scheduling test, coverage feedback, and vulnerability detection. The specific steps are as follows.

[0060] 1. In the static analysis stage, the present invention identifies the mutation scope in the protocol client, mainly involving core logics such as message construction and encrypted transmission, and filters out functions and data structures suitable for instrumentation. The algorithm for the static analysis stage is as Figure 3 shown.

[0061] (1) By analyzing the source code of the protocol client, identify functions involving system calls (such as write, send, sendto). These functions are usually used for sending protocol messages. For example, in OpenSSL, ssl3_do_write is responsible for sending TLS messages, and tls_construct_client_hello is responsible for constructing TLS handshake messages.

[0062] (2) From the parameters of these message sending functions, extract the session structure (such as SSL_CONNECTION), which contains key information such as protocol version, encryption algorithm and key, certificate information, and handshake status. On this basis, use Clang / LLVM for static analysis to filter all functions that take the session structure as a parameter. These functions usually involve message construction, encryption operations, and parsing operations. For example, tls_construct_finished is responsible for generating the TLS termination message, AES_encrypt performs encryption, and SSL_read is responsible for decrypting the data packet. These key functions constitute the target scope of mutant point instrumentation.

[0063] 2. In the mutant point instrumentation stage, the present invention performs mutant point instrumentation on the data flow and control flow instructions of the protocol client during compilation to ensure that the mutation operation is performed before encryption, thereby generating test cases with valid encryption but mutated data.

[0064] (1) In terms of data flow mutation, the present invention inserts mutation logic at load (loading) and store (storing) instructions. For example, by instrumenting the load instruction, the loaded value can be modified (such as incrementing by 1, setting to 0, a random value) to test the protocol parsing logic; by instrumenting the store instruction, the stored value can be mutated to test the abnormal behaviors of key negotiation and certificate verification.

[0065] (2) In terms of control flow mutation, the present invention inserts mutation logic at control instructions such as cmp (comparison) and switch (branch). For example, flipping the cmp comparison result can change the protocol logic judgment, causing the client to skip the CertificateVerify message, thus simulating the authentication bypass vulnerability (CVE-2022-25640); modifying the branches of the switch statement can adjust the protocol state jump and trigger unexpected behaviors.

[0066] (3) In the specific implementation, the present invention uses LLVM for instrumentation, automatically inserts mutant points for load, store, cmp, and switch instructions that meet the mutation scope, and generates a unique mutant point ID to facilitate the subsequent dynamic control of the activation of mutant points by the scheduling end.

[0067] 3. In the scheduling end control stage, the present invention is responsible for managing the activation of mutant points and scheduling the communication between the mutant client and the system under test (SUT). The scheduling algorithm is as Figure 4 shown.

[0068] (1) In terms of mutant point selection, when the scheduling end is initialized, it will traverse all mutant points, execute tests individually, and record the coverage feedback. If a certain mutant point can trigger new coverage, it will be added to the high-priority mutant point queue (PMP). In subsequent tests, the scheduling end will preferentially select PMP for mutation to optimize the test efficiency.

[0069] (2) In the combined mutation strategy, the present invention supports single-point mutation testing and multi-point mutation testing. Single-point mutation testing means that only one mutant point is activated each time to observe its impact on the protocol execution process; multi-point mutation testing is to activate multiple mutant points simultaneously to explore the complex interaction states of the protocol, such as error state transition and abnormal authentication process.

[0070] (3) In terms of interaction management, the present invention ensures that the interaction between the mutant client and the server is executed independently to prevent state pollution, and monitors abnormal responses, such as unexpected error codes, process crashes, or abnormal state machine transitions, so as to identify possible vulnerabilities.

[0071] 4. In the coverage feedback and vulnerability detection stage, the present invention guides subsequent tests through coverage analysis.

[0072] (1) The system will record the baseline coverage of the non-mutated client. Then, in the mutation test, by comparing the coverage, if it is found that the mutated test case triggers a new code path, the mutant point will be added to the priority mutation queue and further in-depth testing will be carried out. If a certain mutant point triggers an abnormal branch, there may be a logic vulnerability.

[0073] (2) The present invention also integrates an automated vulnerability analysis module, which can detect abnormal behaviors, such as server crashes, error return values, or abnormal protocol states, and further analyze the exploitable nature of the stably reproducible abnormal behaviors.

[0074] 5. In the vulnerability analysis and reporting stage, the present invention will generate a complete vulnerability report, including the test cases that trigger the vulnerability, mutant point information (including data comparison before and after mutation), vulnerability impact analysis, and CVE application suggestions. In practical applications, the method of the present invention has successfully discovered multiple protocol implementation vulnerabilities. The test results show that this method can significantly improve the protocol state coverage and code branch coverage. In the TLS1.3 protocol test, the protocol state coverage is increased by 44.69% (as shown in Table 1), and the code branch coverage is increased by 16.36% (as Figure 5 shown), and 4 new vulnerabilities are successfully discovered.

[0076]

[0077] Table 1 Comparison of protocol states covered by the fuzzing method based on program mutation and the relatively better-performing method SGFuzz in current protocol fuzzing in protocol fuzzing

[0078] The present invention proposes a protocol fuzzing method based on instruction-level mutation. Through static analysis, compile-time instrumentation, scheduling tests, and coverage feedback, it effectively solves the problem of protocol state reachability and proves its efficiency and generality in actual tests. Experiments show that this method has discovered multiple new vulnerabilities in 6 protocol implementations such as OpenSSL and wolfSSL, and can be widely applied to vulnerability mining of security protocols such as TLS, SSH, and MQTT.

[0079] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A protocol fuzz testing method based on instruction-level mutation, characterized in that The method includes the following steps: Step 1, Static analysis stage: Determine the mutation range in the protocol client through static analysis, identify functions related to message construction and transmission, and extract the session structure as the core basis for the mutation range; Step 2, Mutation point instrumentation stage: Instrument the data flow and control flow instruction operations in the client program during compilation to generate mutation points, ensure that the mutation operations are performed before encryption, and generate test cases with valid encryption but mutated data; Step 3, Scheduling fuzz testing stage: Manage the communication between the mutated client and the system under test (SUT) through the scheduler, use the coverage feedback to guide the selection of mutation points, where the coverage feedback includes the coverage of the state enumeration values defined by the protocol state variables, and explore deeper protocol states by increasing the coverage count of the state enumeration values to discover potential vulnerabilities.

2. The protocol fuzz testing method based on instruction-level mutation according to claim 1, wherein Step 1 is specifically as follows: 1-1. Identify the message sending function in the client through static analysis, and extract the session structure in its parameters as the core basis for the mutation range; 1-2. Generate an intermediate representation by compiling the source code of the client program. Based on the function call analysis of this intermediate representation, identify the functions in the client source code that call the system calls write, send, or sendto as message sending functions, and extract the first parameter of the message sending function as the session structure as the core basis for the mutation range.

3. The protocol fuzz testing method based on instruction-level mutation according to claim 1, characterized in that In Step 2, it is specifically as follows: 2-1. Instrument the data flow instructions, including mutating the load and store instructions, to ensure that the mutated data can still pass the verification in the decryption stage after encryption; 2-2. Instrument the control flow instructions, specifically including reversing the result of the compare (cmp) instruction and mutating the variable value loaded by the branch (switch) instruction within the valid case range of this switch to change the program execution path, thereby generating an abnormal protocol message sequence to trigger potential vulnerabilities.

4. The protocol fuzzing test method based on instruction-level mutation according to claim 1, characterized in that Step 3 is specifically as follows: 3-1. The scheduler assigns a unique identifier to each mutation point and controls the activation of the mutation point through a global variable; 3-2. The scheduler first activates all mutation points, screens out the mutation points that can trigger more than the baseline coverage (i.e., the server coverage when not mutated), and defines them as priority mutation points (PMPs). In subsequent fuzz testing, the scheduler preferentially selects mutation points from the set of priority mutation points for activation to improve the fuzz testing efficiency; 3-3. After selecting an initial mutation point, the scheduler randomly selects additional mutation points to activate simultaneously with it with a certain probability to form a combined mutation of multiple mutation points to explore the complex protocol states and abnormal behaviors triggered by multiple mutations.

5. The protocol fuzz testing method based on instruction-level mutation according to claim 2, wherein, Step 1-1 is specifically as follows: 1-1-1. Identify the functions that call specific system calls by analyzing the client source code as message sending functions; 1-1-2. Extract the first parameter of the message sending function as the session structure, which contains the key information required for the protocol connection; Step 1-2 is specifically as follows: 1-2-1. Automatically screen out the functions with session structure as parameters through a static analysis tool (Clang / LLVM) as session functions; 1-2-2. The session functions include message construction, conversion, and transmission functions to ensure that the mutation range covers the key logic of protocol implementation.

6. The protocol fuzz testing method based on instruction-level mutation according to claim 3, wherein Step 2-1 is specifically as follows: 2-1-1. Insert mutation logic after the load instruction and before the store instruction to ensure that the mutation operation takes effect before the data enters the memory; 2-1-2. Avoid mutating the instructions with pointer type operands to prevent program crashes; 2-1-3. Adopt mutation strategies such as zero values, boundary values, and random values to generate test cases that can trigger integer overflows or out-of-bounds accesses; Step 2-2 is specifically as follows: 2-2-1. Mutate the comparison instructions to change the program execution path by reversing the comparison results; 2-2-2. Mutate the branch instructions to generate abnormal protocol sequences by changing the branch conditions and trigger vulnerabilities related to control flow.

7. The protocol fuzzing test method based on instruction-level mutation according to claim 4, characterized in that, Step 3-2 is specifically as follows: 3-2-1. The scheduler first activates all mutation points, records the mutation points that can trigger new coverage as the priority mutation points (PMPs); 3-2-2. In subsequent fuzz testing, preferentially select the priority mutation points for activation to improve testing efficiency; Step 3-3 is specifically as follows: 3-3-1. The scheduler supports activating multiple mutation points simultaneously to simulate complex protocol interaction scenarios; 3-3-2. The scheduler dynamically adjusts the activation combination of mutation points according to the coverage feedback to ensure that the test cases can cover more protocol states.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the protocol fuzz testing method based on instruction-level mutation as described in any one of claims 1 to 7 above.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, it implements the protocol fuzz testing method based on instruction-level mutation as described in any one of claims 1-7.