Fuzzy testing method, device and equipment based on large language model
Through a large language model, RFC standard documents are parsed and test cases that meet the protocol characteristics are generated, which solves the efficiency and accuracy problems of AFLnet in complex network protocol testing, and achieves more efficient protocol verification and security improvement.
Patent Information
- Application Number
- CN202510178637.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-01
AI Technical Summary
When existing AFLnet fuzz testing tools deal with complex network protocols such as IPV6, TLS, and DNS, it is difficult to generate highly targeted test cases, and have limited verification capabilities for protocol-level security attributes.
The large language model is used to analyze and parse RFC standard documents, extract protocol rules and constraints, generate test cases that meet the protocol characteristics, and verify the consistency with RFC standard documents through differential testing.
It improves the accuracy and security of the implementation of complex network protocols, optimizes the testing efficiency and coverage, and significantly improves the semantic correctness and diversity of test cases.
Smart Images

Figure CN120238477A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of network security and software testing, and particularly to a fuzz testing method, apparatus, and device based on a large language model. Background Art
[0002] AFLnet is a network protocol testing tool developed based on the AFL (American Fuzzy Lop) fuzz testing tool. Its core feature lies in the introduction of state-aware fuzz testing and multi-dimensional feedback guidance mechanisms. AFLnet uses a state machine to simulate different states and transitions of network protocols, which enables it to effectively infer protocol states and generate meaningful multi-round interaction sequences. The use of the state machine allows AFLnet to understand and track the execution flow of the protocol, thereby generating test cases that are more in line with the protocol semantics. The feedback mechanism of AFLnet is its innovation. It not only utilizes traditional code coverage information but also introduces state machine construction based on message history. By analyzing the message interaction sequences during the testing process, AFLnet can dynamically construct and update the state machine model of the protocol. AFLnet effectively explores different execution paths of the program and various states of the protocol by continuously selecting, pruning, mutating, and executing test cases, and using these two types of feedback information to guide the testing process, thus discovering potential vulnerabilities.
[0003] However, when AFLnet is applied to complex network protocols such as IPV6, TLS, and DNS, its efficiency is significantly reduced. This is mainly due to several aspects: First, the state transitions of many modern network protocols are very complex. For example, TLS involves key exchange and session recovery, IPV6 includes address auto-configuration and neighbor discovery, and DNS needs to handle recursive queries and caching mechanisms, etc. These special mechanisms all increase the difficulty of state inference. Although the state machine construction mechanism of AFLnet is very effective in dealing with simple network protocols, when faced with these complex protocols, it is difficult to accurately capture the implicit states and complex operation sequences in the protocol. Therefore, when dealing with modern network protocols, AFLnet is difficult to generate highly targeted test cases, and it mainly focuses on discovering implementation vulnerabilities such as memory errors, and has limited ability to verify protocol-level security attributes. Summary of the Invention
[0004] This application provides a fuzz testing method, apparatus, and device based on a large language model to solve the problems in the above background art.
[0005] In a first aspect, this application provides a fuzz testing method based on a large language model, including:
[0006] Analyze and parse the RFC standard documents of various network protocols using a large language model, and extract protocol rules and constraint conditions;
[0007] Generate test cases that conform to the protocol characteristics using a large language model-driven approach according to the protocol rules and constraints;
[0008] Perform differential testing through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
[0009] Further, before analyzing and parsing the RFC standard documents of various network protocols using a large language model to extract protocol rules and constraints, it also includes:
[0010] Obtain the RFC standard documents of various network protocols and remove the non-core content in the documents;
[0011] Extract the hierarchical structure of the document by identifying the standardized section numbering and title format, and reorganize the content while maintaining semantic integrity.
[0012] Further, before analyzing and parsing the RFC standard documents of various network protocols using a large language model to extract protocol rules and constraints, it also includes:
[0013] Chunk the RFC standard document and ensure the semantic coherence and context integrity of the content after chunking, and assign structured meta-information tags to each text chunk.
[0014] Further, analyzing and parsing the RFC standard documents of various network protocols using a large language model to extract protocol rules and constraints includes:
[0015] Use the prompt architecture of the large language model to identify, analyze, and parse the RFC standard document according to preset rules to extract protocol rules and constraints.
[0016] Further, after analyzing and parsing the RFC standard documents of various network protocols using the prompt architecture of the large language model to extract protocol rules and constraints according to preset rules, it also includes:
[0017] Associate and integrate the rules from different text chunks to obtain a structured rule set while maintaining the semantic consistency of the original text.
[0018] Further, generating test cases that conform to the protocol characteristics using a large language model-driven approach according to the protocol rules and constraints includes:
[0019] Define a policy description language, which includes test actions, location information, and relative positions;
[0020] Generate test cases that include complete test strategies and expected feedback according to the protocol rules and constraints.
[0021] Further, the differential testing is performed through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document, including:
[0022] Send the abnormal packets generated by the test strategy to the target server, collect the actual responses returned by the server and the expected responses predefined by the test strategy when generating the abnormal packets;
[0023] If the actual response is exactly the same as the expected response, directly proceed to execute the next test case;
[0024] If there are differences between the actual response and the expected response, it is determined that there are potential non-compliance points in the software implementation of the network protocol, which are handed over to manual review for judgment, a detailed test report is generated, and the report is sent to the relevant developers of the network protocol software for repair.
[0025] In a second aspect, the present application provides a fuzz testing device based on a large language model, including:
[0026] A rule extraction module, configured to analyze and parse the RFC standard documents of various network protocols by using a large language model, and extract protocol rules and constraint conditions;
[0027] A test case generation module, configured to generate test cases that conform to the protocol characteristics by using a method driven by a large language model according to the protocol rules and constraint conditions;
[0028] A fuzz testing module, configured to perform differential testing through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
[0029] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the fuzz testing method based on a large language model as described above is implemented.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the fuzz testing method based on a large language model as described above is implemented.
[0031] The above technical solutions of the present application have the following advantages:
[0032] The fuzz testing method based on large language models provided in the first aspect of this application analyzes and parses RFC standard documents of various network protocols by using large language models, extracts protocol rules and constraints, and provides richer and more accurate protocol semantic information for the testing process. Then, according to the protocol rules and constraints, a large language model-driven method is used to generate test cases that conform to the protocol characteristics, enhancing the semantic correctness and diversity of the test cases, making the generated test cases more in line with the complex state transitions and field dependencies of various network protocols, and at the same time being able to explore more potential boundary cases and abnormal scenarios. Finally, differential testing is performed through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document, improving the accuracy and security of the complex network protocol implementation, and at the same time optimizing the test efficiency and coverage.
[0033] It can be understood that the beneficial effects of the above-mentioned second aspect, third aspect and fourth aspect can refer to the relevant descriptions in the above-mentioned first aspect, and will not be elaborated here. Description of the Drawings
[0034] In order to more clearly illustrate the specific implementation manners of this application or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific implementation manners or the prior art. Obviously, the following drawings are some implementation manners of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is the flowchart of the fuzz testing method based on large language models provided by this application;
[0036] Figure 2 It is the execution flowchart of RFCDiff provided by this application;
[0037] Figure 3 It is the flowchart of extracting RFC8446 rules based on the rule extraction prompt word LLM provided by this application;
[0038] Figure 4 It is the flowchart of parsing the ClientHello message structure of the TLS1.3 protocol provided by this application;
[0039] Figure 5 It is the process diagram of collecting potential abnormal messages provided by this application;
[0040] Figure 6 It is the structure diagram of the fuzz testing device based on large language models provided by this application;
[0041] Figure 7 It is the structure diagram of the electronic device provided by this application. Detailed Implementation Manner
[0042] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0043] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0044] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.
[0045] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "Multiple" means "two or more".
[0046] The present application proposes a fuzz testing method based on a large language model, which integrates large language model technology and semantic rule analysis of RFC standard documents. This method performs differential testing by comparing the consistency between the actual software implementation of the target protocol and the rule description in the RFC standard document, aiming to improve the accuracy and security of complex network protocol implementations, while optimizing test efficiency and coverage.
[0047] The following will further describe in detail the specific implementation manner of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application but are not used to limit the scope of the present application.
[0048] As Figure 1As shown below, the embodiment of the present application provides a fuzz testing method based on a large language model, which specifically includes the following steps: Analyze and parse the RFC standard documents of various network protocols using a large language model, and extract protocol rules and constraints; According to the protocol rules and constraints, use a large language model-driven method to generate test cases that conform to the protocol characteristics; Perform differential testing through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
[0049] To address the limitations of AFLnet in complex network protocol testing, the embodiment of the present application proposes the RFCDiff tool. RFCDiff introduces several innovative improvements on the basis of AFLnet. First, it uses a large language model (LLM) to automatically analyze and parse the RFC documents of various network protocols, extracting precise protocol rules and constraints, which provides richer and more accurate protocol semantic information for the testing process. Second, RFCDiff adopts a large language model-driven intelligent method in both constraint design and mutation strategies, and can generate test cases that are more in line with the specific protocol characteristics. RFCDiff introduces a state machine constraint violation generation mechanism, which is not only used to identify potential protocol-level logic vulnerabilities, but more importantly, it is used as a guide for generating mutation operators. By analyzing possible state machine constraint violations, RFCDiff can generate more intelligent and targeted mutation operators.
[0050] As Figure 2 shown, for the execution flow of RFCDiff, compared with AFLnet, RFCDiff is based on an agent enhanced by prompt engineering to achieve automatic extraction of RFC rules, automatic parsing, and automatic generation of test strategies. In the fuzz testing scenario for network protocol packets, combined with the RFC rule description, it takes into account the diversity and semantic correctness of mutated packets, effectively improving the fuzz testing efficiency of the tool. This method significantly enhances the semantic correctness and diversity of test cases, making the generated test cases more in line with the complex state transitions and field dependencies of various network protocols, and at the same time can explore more potential boundary cases and abnormal scenarios. This tool has been used to test multiple important network protocol implementations, including Schannel.dll (TLS implementation) on Windows, the IPV6 stack in network devices, and mainstream DNS server software. The test results show that compared with AFLnet, RFCDiff has achieved significant improvements in code coverage and vulnerability discovery capabilities.
[0051] In some embodiments, before using a large language model to analyze and parse RFC standard documents of various network protocols and extract protocol rules and constraints, the following steps are also included: obtaining RFC standard documents of various network protocols and removing non-core content from the documents; extracting the hierarchical structure of the documents by identifying standardized section numbers and title formats, and reorganizing the content while maintaining semantic integrity.
[0052] In the embodiments of the present application, the standard documents of the target protocol are first obtained from authoritative sources, including but not limited to RFC documents of core network protocols such as TLS, IPV6, DNS, etc. For protocols with multiple related standard documents, the main protocol document and its extension documents are supported to be processed simultaneously, so as to ensure that all aspects of the protocol specifications are captured completely. In the document cleaning stage, intelligent text processing technologies are adopted to systematically remove non-core content from the documents. This includes deleting repeatedly appearing header and footer information, cleaning status descriptions, copyright statements, abstracts, tables of contents, etc. at the beginning of the document, and removing auxiliary information such as references and appendices at the end of the document. This processing method not only improves the information density of the document, but also significantly reduces the noise interference in the subsequent analysis process.
[0053] The embodiments of the present application adopt a section structure processing method based on semantic understanding. This method accurately extracts the hierarchical structure of the document by identifying standardized section numbers and title formats, and reorganizes the content while maintaining semantic integrity. In this process, the embodiments of the present application pay special attention to maintaining the expression form of key information in the protocol specifications, including the structural integrity of important contents such as message format definitions, state transition rules, field constraints, etc.
[0054] In some embodiments, before using a large language model to analyze and parse RFC standard documents of various network protocols and extract protocol rules and constraints, the following steps are also included: splitting the RFC standard documents and ensuring the semantic coherence and context integrity of the content after splitting, and assigning structured meta-information tags to each text block.
[0055] To adapt to the processing characteristics of large language models, the embodiments of the present application design a document splitting strategy. This strategy takes into account the input limitations of the model while focusing on ensuring the semantic coherence and context integrity of the content after splitting. Each text block is assigned structured meta-information tags, including protocol identification, version information, section relationships, etc. These information provide important context support for subsequent intelligent analysis. Finally, an efficient document management system is established to achieve structured storage and rapid retrieval of the processed documents. This storage system not only supports classification management according to protocol types and versions, but also maintains complex section hierarchical relationships and association information between documents, providing a reliable data basis for protocol analysis and verification work.
[0056] Table 1 shows the description of the Ticket Age constraint of the PSK extension in the handshake protocol extension of RFC8446. This section details the client's requirements for the ticket expiration, including the timing starting point, usage restrictions, and obfuscation processing mechanism.
[0057]
[0058] In some embodiments, the use of a large language model to analyze and parse RFC standard documents of various network protocols to extract protocol rules and constraints includes: using a prompt word architecture of a large language model to analyze and parse RFC standard documents according to preset rule recognition standards to extract protocol rules and constraints.
[0059] In the design of prompt words for large language models, a well-structured prompt word framework consists of three core components: system prompts, example prompts, and main prompts. These three parts are organically combined to build a complete prompt system.
[0060] As the basic setting of the framework, the system prompt establishes a clear role positioning and behavioral guidelines for the model. It needs to concisely and clearly explain the model's position in the conversation and set the corresponding behavioral boundaries. This global setting ensures that the model can maintain consistent professional performance throughout the interaction process and will not deviate from the expected goals.
[0061] Example prompts guide the model to understand the task requirements through specific examples. By showing standard input-output pairs, the model can accurately grasp the form and quality standards of the expected output. This example-based learning method can effectively reduce the model's understanding bias and improve the accuracy of the output. Example prompts actually provide a reference standard for the model to better understand the task objectives.
[0062] The main prompts use a structured expression between natural language and programming language. It organizes the content through a predefined keyword system (such as @instruction, @command, @rule, etc.), and uses grammatical elements such as indentation and brackets to express hierarchical relationships. This structured design divides the prompt content into different functional modules, such as the command part, the rule part, the format part, and the example part, each of which has a specific task guidance responsibility. At the same time, by clearly defining the boundary conditions and execution rules of various operations, ambiguous understanding of the model is avoided. This design method not only maintains the flexibility of the prompts, but also reduces the human expression burden and the model's cognitive burden through standardized expressions, thereby improving the reliability and consistency of the prompt effect.
[0063] Through this three - layer structure of the prompt framework, it is possible to effectively guide the model to understand the task requirements, ensure the output quality while maintaining the consistency and reliability of the generated results. This structured prompt design method not only improves the model's understanding accuracy but also provides a clear guiding direction for the creation of prompts.
[0064] To achieve the precise extraction of rules in RFC documents, the embodiments of this application design a system prompt architecture for the GPT4 model. Based on the professional perspective of protocol security experts, this architecture systematically integrates the RFC document specification definition and rule recognition methods into the prompts. The prompt structure covers core elements such as complete persona definitions, professional term explanations, context control, and instruction requirements, and demonstrates the standardized input - output format requirements through specific example data.
[0065] In the specific implementation process, a standardized input template is designed based on the constructed prompt system. As shown in Table 2, the prompt for this agent system requires taking text fragments of RFC documents as input and guiding the model to conduct systematic analysis and rule extraction of protocol specifications through preset format specifications. Practice shows that this template can effectively handle complex protocol specification content including client - side behavior permissions and server - side mandatory requirements. When the text is input into the prompt system, the model can generate structured rule description outputs, which accurately capture the specification requirements in the original protocol document and present key rules such as client - side data flow processing and server - side response timing in a clear format.
[0066] The core of this agent system is to construct an agent with the role positioning of a protocol expert through the provided professional prompts. This agent is endowed with the professional knowledge background of deeply understanding network security protocols, covering various protocols such as HTTP, IP, ARP, TLS, etc. Through a carefully designed term system and rule recognition instructions, the agent can accurately understand the professional terms and specification requirements in RFC documents. During the actual operation process, the system inputs the pre - processed text blocks into the agent one by one, and the agent analyzes and extracts them according to the preset rule recognition criteria.
[0067] In the rule recognition stage, the agent focuses on the key language features in the text. First is the recognition of modal verbs (such as must, should, etc.), which usually identify the mandatory or recommended requirements for protocol implementation. Second is the analysis of conditional constraint words (such as if, when, etc.), which often indicate the protocol behavior specifications under specific conditions. The advantage of the agent is that it can not only recognize single rule statements but also understand complex protocol semantics, especially being good at dealing with those specification descriptions that contain both sender and receiver requirements.
[0068]
[0069]
[0070]
[0071] In some embodiments, after using the prompt architecture of the large language model to identify, analyze, and parse the RFC standard document according to preset rules, and extract protocol rules and constraints, it further includes: associating and integrating the rules from different text blocks to obtain a structured rule set while maintaining the semantic consistency of the original text.
[0072] To ensure the accuracy and practicality of the extraction results, a systematic result integration is performed after rule extraction. This integration process organically organizes and associates the rules from different text blocks to ensure that the final output rule set has good structural characteristics while maintaining the semantic consistency of the original text. This structured output form greatly facilitates the subsequent application and analysis of the rules. Table 3 shows a sample input of the rule extraction model, which is a natural language description passage containing potential rules.
[0073]
[0074] The content of Table 4 is the output obtained after inputting Table 3 into the agent optimized by the prompt in Table 2.
[0075] A total of 2 rules are obtained, and each rule has a mandatory or required description.
[0076]
[0077] The structuring of RFC document rules is a method of converting the specification requirements in RFC documents into structured rules, which is mainly achieved through a carefully designed large language model prompt. In the design of the prompt, a multi-level structured framework is adopted. The model is positioned as a protocol security expert through @Persona, and the definitions of key terms such as RFC, specification requirements, roles, and rule types are clarified through @Terminology. To ensure the accuracy of rule generation, strict control conditions are set in the @ContextControl part of the framework, and @Instruction details the specific steps and requirements for rule generation.
[0078] In practical applications, the input received by the model is a natural language description containing multiple elements, including the RFC document name, chapter title, context information, and specific specification requirements. Based on this input information, the model will, through analysis and understanding, convert the unstructured text into a structured rule representation. For example, in the given example, the model received the specification description of the legacy_version field in the ClientHello message of the TLS 1.3 protocol and converted it into a structured rule form.
[0079] Through this structured conversion process, the specification requirements originally scattered in the RFC document are transformed into a set of clear and executable rules. This not only improves the accuracy of protocol implementation but also provides a basis for the automated verification and testing of the protocol. The value of this method lies in its ability to convert the protocol requirements described in natural language into a machine-processable structured form, thereby reducing ambiguity and errors in the protocol implementation process.
[0080]
[0081]
[0082]
[0083] As shown in Table 6, the model input contains 4 attributes, namely RFC document name (RFC Name), chapter name (Chapter Title), passage content (Content), and single potential rule description (SpecificationRequirement).
[0084]
[0085]
[0086] As shown in Table 7, the converted output adopts the JSON format and contains key attributes such as message type, related fields, construction rules, and processing rules. Among them, the construction rules and processing rules are a pair of complementary rules, which respectively describe the behavior requirements of the message sender and receiver. The construction rules clearly stipulate the specific requirements during message construction, while the processing rules detail the verification and processing procedures of the receiver after receiving the message. This paired rule structure ensures the integrity and reliability of protocol communication.
[0087]
[0088] Figure 3It is the algorithmic process from the RFC document to the rule set. The PreProcess function is the preprocessing function for the RFC document, which is used to delete the content irrelevant to the core of the specification. Then, traverse the remaining set of document fragments, call the SemanticAnalysis function, take the document fragments as input, and pass them into the large language model, and output a set of natural language descriptions that can be regarded as rules. Then, traverse the set of natural language descriptions, call the ExtractRule function, and parse the natural language descriptions regarded as rules into a rule set containing multiple attributes.
[0089] In some embodiments, the method for generating test cases that conform to the protocol characteristics by using the large language model-driven method according to the protocol rules and constraints includes: defining a policy description language, where the policy description language includes test actions, location information, and relative positions; generating test cases that include complete test policies and expected feedback according to the protocol rules and constraints.
[0090] In the test environment setup stage, a virtual machine configured with the Windows Server 2019 operating system needs to be prepared. It is recommended that the virtual machine be configured with at least a 2-core processor, 4GB of memory, and 50GB of hard disk space to ensure the smooth operation of the test environment. In terms of the software environment, Internet Information Services (IIS) 10.0 needs to be installed as the web server, and Wireshark / tshark version 3.6.x and above needs to be deployed for network packet analysis, and a man-in-the-middle proxy software such as Fiddler or Burp Suite needs to be prepared for SSL traffic interception.
[0091] After completing the basic environment setup, the IIS server needs to be configured. Add the web server role through the server manager and install the necessary IIS components including default documents, HTTP errors, static content, and SSL certificate services. Subsequently, perform the SSL certificate configuration. First, generate a self-signed SSL certificate, bind the certificate to the HTTPS service in the IIS manager, and configure port 443 for secure communication listening. This step ensures the basic environment for TLS encrypted communication.
[0092] Next, configure the man-in-the-middle proxy server, set the local listening port to 8080, and point the upstream proxy to the already configured IIS server. After enabling the SSL interception function, the root certificate of the proxy server needs to be installed in the test system and the certificate trust chain needs to be configured to ensure the normal progress of the certificate verification process. Verify the availability of the certificate configuration by accessing the test page to ensure that the man-in-the-middle proxy can intercept SSL / TLS traffic normally.
[0093] Such as Figure 4As shown, in the packet capture phase, the tshark tool is used to capture network traffic. By executing the command "tshark -i 1 -f \"tcp port 443\" -w capture.pcap" in the command line, the TCP traffic on port 443 can be captured and saved as a pcap file. To focus on the ClientHello packet during the TLS handshake process, the command "tshark -r capture.pcap -V -Y \"ssl.handshake.type==1\"" is used for packet filtering and parsing. During the parsing process, key information such as the protocol version number, cipher suite list, random number, session ID, and extension fields is focused on. Table 8 shows the packet structure template obtained by parsing the hijacked ClientHello packet using the tshark tool.
[0094]
[0095]
[0096]
[0097] In terms of the expression of the test strategy, the embodiment of the present application defines a complete set of strategy description languages. It includes elements such as test actions (SET, DELETE, DUPLICATE, INSERT, MODIFY), location information (position), relative position (relative_to), etc., making the generated test strategy not only readable but also maintaining a sufficient degree of formality to support automated execution. This precisely targeted test method is significantly different from the methods of randomly mutating or generating heuristic rules in traditional fuzz testing, not only improving the test efficiency but also reducing the generation of invalid test cases.
[0098] In this stage, when designing the output format, it not only includes specific test operation instructions but also includes the description of the expected result (expected_result). This design enables the test execution system to automatically determine whether the test result meets the expectations, greatly improving the automation degree of the test. Compared with traditional fuzz testing that only focuses on the triggering of abnormal behaviors, this solution provides a test determination criterion through clear expected results, significantly enhancing the verifiability and accuracy of the test results.
[0099] Tables 10 and 11 show examples of this stage. For the test of the key_share field in the ClientHello message of the TLS 1.3 protocol, the system can automatically generate test cases containing complete test strategies and expected feedback according to the input rules. This automated strategy generation method effectively reduces the labor cost of protocol test case design while improving the standardization and integrity of testing. Due to the adoption of a standardized strategy description format and clear definition of test actions, the repeatable execution of test cases is ensured, which is a characteristic difficult to guarantee by traditional fuzz testing.
[0100] The design of this stage fully considers scalability. By defining a unified input / output interface format, the system can easily adapt to different types of rule inputs and generate corresponding test strategies, with good versatility. This flexible design enables the system to not require a re-design of test strategies when facing new protocol features, significantly reducing the test maintenance cost. At the same time, since the generated test strategies are all valid test scenarios derived from specific rules, the utilization efficiency of test resources is greatly improved.
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107] In some embodiments, the differential testing is performed through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document, including: sending the abnormal message generated by the test strategy to the target server, collecting the actual response returned by the server and the expected response predefined by the test strategy when generating the abnormal message; if the actual response is exactly the same as the expected response, directly proceed to the execution of the next test case; if there is a difference between the actual response and the expected response, it is determined that there is a potential non-compliance point in the network protocol software implementation, which is handed over to manual review for judgment, and a detailed test report (including test cases, test software versions, expected feedback, actual feedback, potential cause analysis, etc.) is generated and sent to the relevant developers of the network protocol software for repair.
[0108] In network protocol fuzz testing, the message field-level mutation operation is the core testing method refined from the test strategy. These mutation operations mainly include three basic operations: add, remove, and update, which respectively correspond to different test scenario requirements.
[0109] The add operation originates from the test strategy for protocol fault tolerance and extensibility. By inserting additional fields or parameters into the original message, this operation simulates scenarios of protocol extension or abnormal data injection. For example, adding undefined extension fields in the TLS handshake message to test the server's handling ability for non-standard extensions; or inserting redundant fields in the HTTP request header to verify the server's parsing fault tolerance.
[0110] The remove operation is refined based on the protocol necessity test strategy. By selectively removing specific fields from the message, it can verify the degree of dependence of the protocol implementation on essential fields. For instance, removing the cipher suite list in the TLS ClientHello message to test the server's response to missing key fields; or removing the required headers in the HTTP request to examine the server's error handling mechanism.
[0111] The update operation comes from the protocol boundary value and abnormal value test strategy. By modifying the content of existing fields, while keeping the message structure intact, it probes the protocol implementation's handling ability for various input values. Specifically, it includes boundary testing of field lengths, replacing field content with special characters, or modifying the data type of field values, etc. This kind of mutation can effectively discover security issues such as overflow vulnerabilities and injection defects.
[0112] These three mutation operations cooperate with each other to jointly form a complete protocol fuzz testing framework. They are not randomly generated, but are carefully designed according to specific test objectives and strategies, and can systematically verify the robustness, security, and standard compliance of the protocol implementation. Through this structured mutation method, potential protocol implementation defects can be discovered more efficiently.
[0113] The test execution and feedback analysis of abnormal messages is an iterative process that optimizes the test case pool through a systematic comparison and screening mechanism. This process mainly includes three key links: message sending, response collection, and result analysis.
[0114] After the abnormal message generated by the test strategy is sent to the target server, the system will collect two types of key information simultaneously: the actual response message or response status code returned by the server, and the expected response predefined by the test strategy when generating the abnormal message. These two types of information constitute the basic data for result analysis. The system will then perform an exact comparison of these two types of information and determine the subsequent processing flow based on the comparison result.
[0115] As Figure 5 shown, in the comparison and analysis phase, if the actual response is exactly the same as the expected response, it indicates that the current test case fails to discover new system behavior characteristics, and the system will directly proceed to the execution of the next test case. However, when there is a difference between the actual response and the expected response, this difference marks the possible discovery of a new behavior pattern of the system. Such test cases with special value will be included in the interest seed set and used as an important input source for the next round of fuzz testing. Through this dynamic feedback collection and analysis mechanism, the system can continuously optimize the test case pool, concentrate test resources on scenarios that are more likely to discover new system behavior characteristics, and thus improve the overall test efficiency and coverage.
[0116] The fuzz testing method based on large language models provided by the embodiments of this application is used to verify the consistency between the actual software implementation of modern network protocols (such as TLS, IPV6, DNS, etc.) and the rule descriptions in RFC standard documents, and proposes the RFCDiff tool. RFCDiff has three key innovations. First, the tool adopts intelligent document parsing technology, uses large language models (LLMs) to automatically analyze and parse RFC documents of various network protocols, and accurately extracts protocol rules and constraint conditions, providing rich and accurate protocol semantic information for the test process. This technology can be applied to the analysis of multiple protocol standards such as TLS, IPV6, and DNS at the same time. Second, RFCDiff adopts an LLM-driven test strategy, using intelligent methods in constraint design and mutation strategies to adaptively generate test cases according to different protocol characteristics. This method significantly improves the semantic correctness and diversity of test samples and can well adapt to the special requirements of different protocols, whether it is the encryption mechanism of TLS, the address configuration of IPV6, or complex scenarios such as DNS query resolution. Finally, RFCDiff innovatively introduces a state machine constraint violation generation mechanism, which is not only used to identify protocol-level logical vulnerabilities, but also serves as an intelligent guidance system for generating mutation operators. By analyzing potential state machine constraint violations, the system can generate more intelligent and targeted mutation operators to ensure that test cases conform to the complex state transitions and field dependencies of various protocols, and at the same time effectively explore boundary cases and abnormal scenarios.
[0117] The embodiments of this application can improve the consistency between network protocol implementation and the rule descriptions in RFC standard documents. When processing handshake processes and encryption mechanisms such as the TLS protocol, it can accurately verify whether the implementation complies with the standard specifications. At the same time, for the complex address configuration and routing selection mechanisms in the IPV6 protocol, as well as the query resolution and cache processing processes in the DNS protocol, it can effectively ensure the accuracy of their implementation, thereby guaranteeing the reliability and standard compliance of various network protocols in practical applications.
[0118] The embodiments of this application can also enhance the security of protocol implementation by discovering potential security vulnerabilities through comprehensive differential testing. Special attention is paid to the key key exchange and authentication mechanisms in encryption protocols to ensure the security of their implementation. At the same time, for the complex state transitions and session management processes in network protocols, as well as various boundary conditions and abnormal situations that may occur during protocol interactions, in-depth security analysis and verification can be carried out to effectively improve the overall security of protocol implementation.
[0119] The embodiments of this application can also optimize the efficiency of protocol testing. By using large language model technology, it can automatically analyze the semantic rules of RFC documents of various network protocols, accurately understand and extract the core requirements in the standard specifications. Based on these understandings, test cases that conform to the characteristics of specific protocols can be intelligently generated, and the deviations between protocol implementation and standard specifications can be quickly identified, significantly improving the efficiency and accuracy of the testing process.
[0120] The embodiments of this application can also improve the test coverage rate by systematically comparing the implementation with the standard alignment. Deeply verify the state transition and data processing logic of the protocol to ensure the integrity and depth of the test. At the same time, it also focuses on testing the adaptability and robustness of the protocol in various network environments, including performance performance and exception handling capabilities under different network conditions, so as to ensure the reliability and stability of protocol implementation in the actual application environment.
[0121] Corresponding to the fuzz testing method based on large language models described in the above embodiments, as Figure 6 shown, the embodiments of this application also provide a fuzz testing device based on large language models. The fuzz testing device based on large language models includes:
[0122] A rule extraction module, which is used to analyze and parse RFC standard documents of various network protocols by using a large language model, and extract protocol rules and constraint conditions;
[0123] A test case generation module, which is used to generate test cases that conform to the protocol characteristics by using a method driven by a large language model according to the protocol rules and constraint conditions;
[0124] A fuzz testing module for performing differential testing through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
[0125] It should be noted that for the information interaction, execution process, etc. between the above-mentioned modules / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0126] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0127] The embodiment of the present application also provides an electronic device, such as Figure 7 shown, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the fuzz testing method based on the large language model provided in the first aspect.
[0128] In applications, the electronic device may include, but is not limited to, a processor and a memory. Figure 7 This is only an example of the electronic device and does not limit the electronic device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, input / output devices, network access devices, etc. The input / output devices may include cameras, audio acquisition / playback devices, displays, etc. The network access device may include a network module for performing wireless network with external devices.
[0129] In an application, the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0130] In an application, in some embodiments, the memory may be an internal storage unit of an electronic device, such as the hard disk or memory of the electronic device. In other embodiments, the memory may also be an external storage device of the electronic device, for example, a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory may also include both the internal storage unit and the external storage device of the electronic device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of a computer program. The memory may also be used to temporarily store data that has been output or will be output.
[0131] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0132] To implement all or part of the processes in the above method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0133] Those of ordinary skill in the art will appreciate that the device and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0134] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. Additionally, the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the devices can be indirectly coupled or communication-connected, which can be in electrical, mechanical, or other forms.
[0135] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A fuzzy testing method based on a large language model, characterized in that: include: Use large language models to analyze and parse RFC standard documents of various network protocols and extract protocol rules and constraints; According to the protocol rules and constraints, a large language model driven method is used to generate test cases that meet the protocol characteristics; Differential testing is performed through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
2. The fuzzy testing method based on a large language model as claimed in claim 1, characterized in that: The use of a large language model to analyze and parse RFC standard documents of various network protocols and extract protocol rules and constraints also includes: Obtain RFC standard documents for various network protocols and remove non-core content from the documents; The hierarchical structure of the document is extracted by recognizing standardized chapter numbering and heading formats, and the content is reorganized while maintaining semantic integrity.
3. The fuzzy testing method based on a large language model as claimed in claim 1, characterized in that: The use of a large language model to analyze and parse RFC standard documents of various network protocols and extract protocol rules and constraints also includes: The RFC standard document is divided into blocks and the semantic coherence and contextual integrity of the content after block division are guaranteed, and each text block is given a structured meta-information tag.
4. The fuzzy testing method based on a large language model as claimed in claim 3, characterized in that: The large language model is used to analyze and parse RFC standard documents of various network protocols to extract protocol rules and constraints, including: Utilize the prompt word architecture of the large language model to analyze and parse the RFC standard document according to the preset rule recognition standards, and extract the protocol rules and constraints.
5. The fuzzy testing method based on a large language model as claimed in claim 4, characterized in that: The method utilizes the prompt word architecture of the large language model to analyze and parse the RFC standard document according to the preset rule recognition standard, and extracts the protocol rules and constraints, and further includes: The rules from different text blocks are associated and integrated to obtain a structured rule set while maintaining the semantic consistency of the original text.
6. The fuzzy testing method based on a large language model as claimed in claim 1, characterized in that: The method of using a large language model driven method to generate test cases that meet the protocol characteristics according to the protocol rules and constraints includes: Defining a policy description language, wherein the policy description language includes test actions, location information, and relative positions; Based on the protocol rules and constraints, test cases are generated that include complete test strategies and expected feedback.
7. The fuzzy testing method based on a large language model as claimed in claim 6, characterized in that: The differential test is performed through the test case to verify the consistency between the actual software implementation of the network protocol and the RFC standard document, including: Sending the abnormal message generated by the test strategy to the target server, collecting the actual response returned by the server and the expected response predefined by the test strategy when generating the abnormal message; If the actual response is completely consistent with the expected response, directly proceed to the execution of the next test case; If the actual response differs from the expected response, it is determined that there is a potential non-compliance in the network protocol software implementation, which is then submitted for manual review and judgment, and a detailed test report is generated and sent to the relevant developers of the network protocol software for repair.
8. A fuzzy testing device based on a large language model, characterized in that: include: The rule extraction module is used to analyze and parse the RFC standard documents of various network protocols using a large language model to extract protocol rules and constraints; A test case generation module, used to generate test cases that meet the protocol characteristics using a large language model driven method according to the protocol rules and constraints; The fuzz testing module is used to perform differential testing through the test cases to verify the consistency between the actual software implementation of the network protocol and the RFC standard document.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the fuzzy testing method based on the large language model as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the fuzzy testing method based on a large language model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
State preprocessing method for network protocol program analysis
CN120639523A
A State Preprocessing Method for Network Protocol Program Analysis
CN120639523B
LLM-based end-to-end industrial control protocol fuzzy test script generation method
CN120768813A
LLM end-to-end-based industrial control protocol fuzz testing script generation method
CN120768813B
Description language tool for protocol fuzz testing
CN120768814A