Some / ip protocol grey box fuzzy test method and system
By identifying the sensitive structure in the on-board Ethernet protocol data packet, using structure-sensitive mutation strategies, using the GRU network and attention mechanism to extract sensitive elements, and building a corpus for fuzz testing, solving the problems of low testing efficiency and low mutation efficiency in the existing technology, and achieving efficient vulnerability detection.
Patent Information
- Application Number
- CN202510484456.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing fuzz testing methods lack a state boot mechanism in the on-board Ethernet protocol, resulting in low testing efficiency and inefficient packet mutation, which cannot effectively improve the mutation efficiency, and it is difficult to meet the strict state sequence and packet structure constraints of the on-board Ethernet protocol.
By identifying sensitive structures in protocol data packets, optimizing mutation strategies, adopting structure-sensitive mutation strategies, improving seed generation quality, using GRU network and attention mechanism to extract sensitive elements, building a corpus for test case generation, and meeting protocol status and format constraints.
Improves the efficiency and vulnerability detection capabilities of the fuzz testing process. The generated test cases can trigger new states and crash locations, comply with the protocol format specifications, reduces the dependence of reverse engineering, and improves the effectiveness of tests.
Smart Images

Figure CN120342689A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a method and system for gray-box fuzz testing of SOME / IP protocol. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the wide application of in-vehicle Ethernet protocol, its security risks also increase accordingly. Attackers can use protocol vulnerabilities to attack in-vehicle systems, affecting the safety and stability of vehicles.
[0004] Currently, fuzz testing, as an efficient vulnerability detection technology, has been widely used in network protocol security testing. Gray-box fuzz testing technologies such as AFLNET, StateAFL, and Ori can effectively detect network protocol vulnerabilities. However, the existing fuzz testing methods still face the following challenges in the security testing of in-vehicle Ethernet protocol:
[0005] (1) Lack of a mechanism for lightweight state-guided fuzz testing: The in-vehicle Ethernet protocol often has strict state sequence rules, and existing methods are difficult to apply to the in-vehicle Ethernet protocol. For example, AFL and Ori lack a state-guided mechanism, resulting in low testing efficiency; AFLNET and StateAFL do not support testing the SOME / IP protocol.
[0006] (2) Inefficient packet mutation: The packets of the in-vehicle Ethernet protocol have strict structure constraints. Traditional fuzz testing tools usually adopt random mutation or equal-probability mutation strategies, which cannot effectively improve the mutation efficiency, resulting in limited testing effects. Summary of the Invention
[0007] To solve the above problems, the present invention proposes a method and system for gray-box fuzz testing of SOME / IP protocol, which automatically identifies sensitive structures in protocol packets to optimize the mutation strategy. Through the structure-sensitive mutation strategy, the quality of seed generation is improved, and the efficiency and vulnerability detection ability of the fuzz testing process are enhanced.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In a first aspect, the present invention provides a method for gray-box fuzz testing of SOME / IP protocol, including:
[0010] Calculate the generation probability for each element of each seed sequence in the obtained protocol data packet, and determine the sensitive element sequence with the element having the maximum generation probability as the sensitive element; and calculate the replication probability based on the seed sequence and the corresponding sensitive element sequence, determine the conditional probability of the sensitive element sequence according to the generation probability and the replication probability, and take maximizing the conditional probability of the sensitive element sequence as the objective function to complete the extraction of sensitive elements;
[0011] Determine the relative offset of the sensitive element with respect to the data packet header of the seed sequence where it is located, determine the ID based on the sensitive element and the relative offset, and determine the weight of the sensitive element according to the generation probability, code coverage rate, and status coverage rate of the sensitive element. Using the ID as an index, construct a corpus with the sensitive element, its relative offset, weight, and response status code;
[0012] For the protocol data packet to be mutated, select the corresponding corpus in the corpus according to the protocol status, and perform replacement mutation with the sensitive element and relative offset corresponding to the maximum weight to obtain a test case, and complete the fuzz testing process by monitoring the execution of the test case.
[0013] As an alternative implementation, the generation probability p c is:
[0014] p c (SS t |SS <t ,S) = o(SS t-1 ,h t ,S′); h t = f(SS t-1 ,h t-t ,S′);
[0015] where h t is the output of the hidden layer of the GRU network at time t, h t-1 is the output of the hidden layer of the GRU network at time t - 1, f(·) is the hidden layer calculation function of the GRU network; SS t is all the sensitive element sequences extracted at time t, SS t-1 is all the sensitive element sequences extracted at time t - 1, o(SS t-1 ,h t ,S′) is the generation probability calculation function, specifically at time t, calculate the generation probability of each element of each seed sequence in the extraction vector S′ as a sensitive element, and take the element with the maximum generation probability as the sensitive element at time t.
[0016] As an alternative implementation, the extraction vector S′ is: after obtaining the protocol data packet, encode each seed sequence in the protocol data packet to obtain the extraction vector, and perform the extraction of sensitive elements based on the extraction vector.
[0017] As an alternative embodiment, the replication probability p s is:
[0018]
[0019] where θ(SS t = S i ) is the output score calculated after the seed sequence S i is replicated to SS t , that is, to determine the element positions of the elements belonging to sensitive elements after the seed sequence S i corresponds to SS t . Thus, the output score is determined through the function θ(SS t = S i ); I is the number of sequences in SS t ; Z is the sum of all scores; hh i is the operation of the output parameter matrix after the representation layer encoding of S i ; ω s is the weight parameter matrix; σ(·) is the tanh non-linear activation function; T is the matrix transpose.
[0020] As an alternative embodiment, the objective function defined to maximize the conditional probability of the sensitive element sequence is:
[0021] P(SS t | SS <t , S) = p c (SS t | SS <t , S) + γp s (SS t | SS <t , S);
[0022] where SS <t is all sensitive element sequences extracted before time t; γ is the harmonic parameter.
[0023] As an alternative embodiment, the weight is: W(ID) = O(·) + Ψ(·) + Φ(·);
[0024] where W is the weight; O(·) is the generation probability of sensitive elements under the current ID index, Ψ(·) is the code coverage rate of sensitive elements under the current ID index, and Φ(·) is the state coverage rate of sensitive elements under the current ID index, which is obtained by dividing the number of relevant states of sensitive elements under the current ID index by the total number of states.
[0025] In a second aspect, the present invention provides a some / ip protocol gray-box fuzz testing system, including:
[0026] The sensitive structure extraction module is configured to calculate and generate probabilities for each element of each seed sequence in the obtained protocol data packet, determine the sensitive element with the highest probability as the sensitive element, and thus determine the sensitive element sequence; calculate the replication probability according to the seed sequence and the corresponding sensitive element sequence, determine the conditional probability of the sensitive element sequence based on the generation probability and the replication probability, and complete the extraction of the sensitive element with the objective function of maximizing the conditional probability of the sensitive element sequence;
[0027] The corpus module is configured to determine the relative offset of the sensitive element with respect to the data packet header of the seed sequence where it is located, determine the ID according to the sensitive element and the relative offset, determine the weight of the sensitive element according to the generation probability, code coverage rate, and status coverage rate of the sensitive element, and construct a corpus with the ID as the index, the sensitive element and its relative offset, weight, and response status code;
[0028] The test module is configured to, for the protocol data packet to be mutated, select the corresponding corpus in the corpus according to the protocol state, and perform replacement mutation with the sensitive element and relative offset corresponding to the maximum weight, thereby obtaining test cases, and complete the fuzz testing process by monitoring the execution of the test cases.
[0029] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0030] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0031] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] When sending data packets during the SOME / IP protocol communication running in a vehicle, it is necessary to simultaneously meet the state sequence constraint and the data packet format constraint. Therefore, when performing fuzz testing on it, two requirements should be met: 1) being able to represent the protocol state relationship and using it to schedule the fuzz testing process; 2) being able to generate test seeds that conform to the protocol syntax format specification. However, existing fuzz testing tools, such as AFL and Ori, lack a state-guided mechanism and have low testing efficiency; AFLNET and StateAFL do not support testing the SOME / IP protocol. To solve the problem of low efficiency in fuzz testing the SOME / IP protocol, the present invention proposes a grey-box fuzz testing method and system for the some / ip protocol, which optimizes the mutation strategy by automatically identifying sensitive structures in protocol data packets. Through the structure-sensitive mutation strategy, the quality of seed generation is improved, and the efficiency of the fuzz testing process and the vulnerability detection ability are enhanced.
[0034] The method of the present invention is aimed at the some / ip protocol, and instead of generating a complete test case, it learns the sensitive elements in the test cases with good performance, enabling the test case to trigger new states, code coverage, the field positions and their characteristics of crashes, and guiding mutations accordingly, so that the obtained seeds are more in line with the protocol format specification, and the utilization of past test experience is also more efficient and fine-grained; it does not require the protocol reverse process and has less restrictions.
[0035] Advantages of additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0037] Figure 1 It is a flowchart of the grey-box fuzz testing method for the some / ip protocol provided in Embodiment 1 of the present invention;
[0038] Figure 2 It is a schematic diagram of the principle of the grey-box fuzz testing method for the some / ip protocol provided in Embodiment 1 of the present invention;
[0039] Figure 3 It is an architecture diagram of the sensitive structure extraction module provided in Embodiment 1 of the present invention;
[0040] Figure 4 It is a schematic diagram of the corpus entry structure provided in Embodiment 1 of the present invention;
[0041] Figure 5 This is the flowchart provided for Embodiment 1 of the present invention;
[0042] Figure 6 This is the first experimental verification result provided for Embodiment 1 of the present invention;
[0043] Figure 7 This is the second experimental verification result provided for Embodiment 1 of the present invention;
[0044] Figure 8 This is the architecture diagram of the some / ip protocol gray-box fuzz testing system provided for Embodiment 2 of the present invention. Detailed implementation manners
[0045] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0046] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0047] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0049] Embodiment 1
[0050] This embodiment provides a some / ip protocol gray-box fuzz testing method, as Figure 1 shown, including:
[0051] Calculate the generation probability for each element of each seed sequence in the obtained protocol data packet, and use the element with the maximum generation probability as the sensitive element, thereby determining the sensitive element sequence; and calculate the replication probability according to the seed sequence and the corresponding sensitive element sequence, determine the conditional probability of the sensitive element sequence according to the generation probability and the replication probability, and use maximizing the conditional probability of the sensitive element sequence as the objective function to complete the extraction of the sensitive element;
[0052] Determine the relative offset of the sensitive element with respect to the data packet header of the seed sequence where it is located. Determine the ID based on the sensitive element and the relative offset. Determine the weight of the sensitive element based on the generation probability, code coverage rate, and status coverage rate of the sensitive element. Using the ID as an index, construct a corpus with the sensitive element, its relative offset, weight, and response status code;
[0053] For the protocol data packet to be mutated, select the corresponding corpus in the corpus according to the protocol status, and perform replacement mutation with the sensitive element and relative offset corresponding to the maximum weight, thereby obtaining test cases. By monitoring the execution of the test cases, complete the fuzz testing process.
[0054] The following combines Figure 2 to elaborate in detail on the method of this embodiment.
[0055] In this embodiment, protocol data packets are extracted from in-vehicle network traffic and used as seed data packets, that is, as the starting point of fuzz testing. Protocol data packets record the real message exchanges between the client and the server, and can be extracted using network sniffers and packet analyzers, such as tcpdump and Wireshark, thus serving as the initial seeds for fuzz testing.
[0056] In this embodiment, based on the sensitive structure extraction module and the corpus module, structure-sensitive mutation is performed on the seed data packet to obtain test cases and send them to the target system.
[0057] As Figure 3 shown, the sensitive structure extraction module includes a Bi-GRU (Bidirectional Gated Recurrent Unit) representation layer with an additional attention mechanism and a sensitive structure extraction layer based on GRU; considering the differences in the formats of protocol data packets, the Bi-GRU representation layer with an additional attention mechanism is used to encode the seed sequence, and the sensitive structure extraction layer based on GRU is used as the decoding network for extracting sensitive elements.
[0058] Specifically:
[0059] (1) The representation layer is built on Bi-GRU to obtain the protocol data packet S=(S1, S2,..., S m ), where the protocol data packet S is a set of unordered and variable-length seed sequences, S i is the i-th seed sequence of S, and n is the total number of sequences; the protocol data packet S is automatically mapped to a high-dimensional embedding space and encoded into an extraction vector S' for the sensitive element extraction task.
[0060] Meanwhile, in natural language processing tasks, the processing of each word is usually the same. However, considering the structural sensitivity of protocol messages, the goal of this embodiment is to capture the important relationships between structures while maintaining good accuracy. Therefore, this embodiment introduces an attention mechanism to enhance the representation layer.
[0061] It can be understood that the Bi-GRU network can implement the above encoding process, which is a conventional method and will not be elaborated specifically.
[0062] (2) Extraction layer, using GRU to decode the high-dimensional extraction vector S′ into a sensitive element sequence SS = (SS1, SS2,..., SS n ), where the i-th sensitive element sequence SS i = (ss i1 , ss i2 ,..., ss ik ), ss ij represents the j-th sensitive element in the i-th sensitive element sequence SS i , and k is the number of sensitive elements in the i-th sensitive element sequence SS i . The sensitive element sequence SS has a variable length, that is, the length of each sensitive element sequence can be different, which means that each seed sequence S i can extract different numbers of sensitive elements.
[0063] To guide the entire extraction process, this embodiment designs an objective function, aiming to maximize the conditional probability of the sensitive element sequence. The objective function is defined as follows to select more sensitive elements in the seed sequence:
[0064] P(SS t |SS <t , S) = p c (SS t |SS <t , S) + γp s (SS t |SS <t , S) (1);
[0065] Among them, SS t is all the sensitive element sequences extracted at time t; P(SS t |SS <t , S) is the conditional probability of the sensitive element sequence SS t at time t; SS <t is all the sensitive element sequences extracted before time t; γ is a harmonic parameter that controls the proportion of the generation probability and the copy probability.
[0066] Maximize the conditional probability of the sensitive element sequence through formula (1) to extract sensitive elements from the original seed sequence. At the same time, by combining structural attention, the position information and field information are encoded together, and the copy probability is added to the objective function, which can not only encode the input content, but also encode the structural information between elements, helping to discover the key parts in the original seed sequence.
[0067] It can be seen from formula (1) that by combining structural attention, the conditional probability p(SS t |SS t ,S) of SS <t not only considers the generation probability p c of generating sensitive elements from the extraction layer, but also includes the copy probability p t of copying from the source input to SS s .
[0068] Specifically: The generation probability p c is:
[0069] p c (SS t |SS <t ,S) = o(SS t-1 ,h t ,S′);
[0070] h t = f(SS t-1 ,h t-1 ,S′);
[0071] Among them, h t is the output of the hidden layer of the GRU network at time t, h t-1 is the output of the hidden layer of the GRU network at time t - 1, f(·) is the calculation function of the hidden layer of the GRU network; SS t is all the sensitive element sequences extracted at time t, SS t-1 is all the sensitive element sequences extracted at time t - 1, o(SS t-1 ,h t ,S′) is the generation probability calculation function, specifically at time t, calculate the generation probability of each element of each seed sequence in the extraction vector S′ being a sensitive element, and take the element with the maximum generation probability as the sensitive element at time t.
[0072] It can be understood that the calculation of the generation probability uses the GRU network. Specifically, the intermediate processing process is the processing of the GRU network, which will not be elaborated, and the output of its last layer is used as the target output.
[0073] The copy probability p s only considers the elements in the source input and is expressed as:
[0074]
[0075] Among them, θ(SS t = S i ) is the output score calculated after copying the seed sequence S i to SS t , that is, to determine the element positions of the sensitive elements after the seed sequence S i corresponds to SS t . Thus, the output score is determined through the function θ(SS t = S i ); I is the number of sequences in SS t ; Z is the sum of all scores, which is used for normalization; hh i is the output parameter matrix operation after the representation layer encoding of S i ; ω s is the weight parameter matrix; σ(·) is the tanh non-linear activation function, which helps to map h t and hh i into the same semantic space; T is the matrix transpose.
[0076] For the sensitive element extraction task, the objective function is to maximize the conditional probability of the sensitive element sequence. Simply put, according to the training batches of the seed sequence S and the corresponding sensitive element sequence SS, the model is trained by minimizing the negative log-likelihood, which is defined as:
[0077]
[0078] Among them, is all the sensitive element sequences extracted at time t in the j-th training batch; is all the sensitive element sequences extracted at time t - 1 in the j-th training batch; S j is the seed sequence of the j-th training batch; N is the total number of training batches; T is the number of time steps.
[0079] In this embodiment, this module adopts a Bi-GRU combined with an attention mechanism to encode the protocol data packet, automatically learn the importance of data elements, enhance the understanding of structural information, and extract key sensitive elements, which will be used to guide mutation during the fuzz testing process to improve the testing efficiency. At the same time, the model is optimized through incremental learning. As the fuzz testing process progresses, the extraction strategy of sensitive structures is dynamically adjusted to improve the recognition accuracy.
[0080] (3) Corpus module: As Figure 4As shown, each corpus entry in the corpus module is represented in the structure of a five-tuple [ID, sensitive element, relative offset, weight, related status]; ensuring the uniqueness of data and preventing redundancy to achieve the weight setting of sensitive elements and the lightweight representation of status relationships.
[0081] Among them, the sensitive element is the output of the sensitive structure extraction module.
[0082] The relative offset represents the byte offset of the sensitive element in the original seed sequence relative to the data packet header of the seed sequence where it is located.
[0083] To ensure uniqueness, the hash values of the sensitive element and the relative offset are calculated as the ID of the corpus entry, which helps to avoid excessive redundancy in the corpus and affect the mutation efficiency.
[0084] The weight represents the possibility of selecting a corpus entry, specifically;
[0085] W(ID) = O(·) + Ψ(·) + Φ(·);
[0086] Among them, W is the weight; O(·) is the generation probability of the sensitive element by the sensitive element extraction module for the sensitive element under the current ID index, Ψ(·) is the code coverage rate of the sensitive element under the current ID index, which is collected by source code instrumentation and is the prior art and will not be elaborated; Φ(·) is the state coverage rate of the sensitive element under the current ID index, which is obtained by dividing the number of related statuses of the sensitive element under the current ID index by the total number of statuses.
[0087] The related status R represents the state range in which the sensitive element under the current ID index runs, which is identified by the response status code (such as 404 Not Found: the URL corresponding resource does not exist (path error or resource deleted), etc.), so as to achieve the lightweight representation of the status relationship to assist in structure-sensitive mutation.
[0088] Among them, the status sequence rule (protocol status relationship) is the protocol state machine and its transition relationship defined when the network protocol is implemented, indicating the working process and response of the network protocol after receiving the seed. The relevant status field of the corpus stores the protocol status information that the corpus can reach, thus lightweightly representing the relationship between states. By retrieving the relevant status, the corpus that causes the state change can be quickly found for mutation.
[0089] When the sensitive structure extraction module generates a new output to be sent to the corpus, the fuzzer calculates the hash values of the sensitive element and the relative offset to check whether the corpus entry already exists through the ID. If it exists, the weight and related status of the corresponding five-tuple are updated; if not, a new five-tuple is created to store the new corpus entry.
[0090] In this embodiment, in the structure-sensitive fuzz testing, the corpus module enhances the effectiveness of mutations. The higher the weight of a corpus entry, the higher its structural sensitivity, and the more likely it is to be selected for seed mutation. When using the corpus to enhance the effect of seed mutation, first, relevant corpus entries are collected according to the current execution state, then the sensitive element corresponding to the maximum weight is selected, and the current seed sequence is replaced and mutated according to the corresponding relative offset. As the fuzz testing progresses, a virtuous cycle of mutation-storage-mutation is established, and the corpus becomes more perfect, which can better guide the fuzz testing.
[0091] In this embodiment, by using the corpus module, instead of mutating each data element equally, sensitive elements are strategically selected from the corpus to replace some elements of the original seed sequence, and this process follows strict data format constraints. In this way, the positions of elements that play a more important role in the fuzz testing process can be automatically identified, and elements with higher value are used for replacement, thereby improving the efficiency of fuzz testing.
[0092] In this embodiment, as Figure 5 shown, after the structure-sensitive seed mutation is completed, the execution of the target system is monitored, and information on crashes, abnormal behaviors, and code coverage is recorded. If the current seed is interesting, that is, it generates a new state, new coverage, or a crash, the sensitive elements are extracted and stored in the corpus to guide subsequent mutations. If the current seed is not interesting, the next round of fuzz testing is performed. The above process is continuously repeated to optimize the seed quality and improve the vulnerability detection ability.
[0093] To evaluate the effectiveness of the method in this embodiment using the structural sensitivity of seeds to guide fuzz testing, based on an implementation example of the some / ip protocol, the number of crashes triggered by the method in this embodiment and the ORI method within three identical time periods was evaluated. The more crashes are triggered, the stronger the vulnerability detection ability of the fuzz testing tool. The detailed results of the experiment are as Figure 6 and Figure 7 shown. The two cases represent different structural sensitivity requirements and are used to more comprehensively evaluate the ability of the method in this embodiment in terms of vulnerability detection.
[0094] Figure 6 represents a simple crash trigger condition, where a crash can be triggered as long as one byte meets the expected structural constraints. Figure 7It represents complex crash-triggering conditions, and crashes can only be triggered when multiple bytes at different positions meet the expected structural constraints. The experimental results show that, compared with Ori, the method of this embodiment achieves the best results in both cases. Among them, in the case of more stringent structural constraints, the method of this embodiment achieves better results, and the number of crashes triggered is 59.18% more than that of Ori. This indicates that the structure-sensitivity-based mutation implemented by the method of this embodiment can better generate seeds that conform to the structural constraints of the in-vehicle Ethernet protocol, thereby being more likely to trigger crashes and achieving more efficient fuzz testing.
[0095] It should be noted that all data is obtained on the basis of compliance with laws, regulations, and user consent, and the data is legally applied.
[0096] Embodiment 2
[0097] As Figure 8 shown, this embodiment provides a some / ip protocol gray-box fuzz testing system, including:
[0098] A sensitive structure extraction module, configured to calculate the generation probability for each element of each seed sequence in the obtained protocol data packet, determine the sensitive element with the highest generation probability as the sensitive element, and thus determine the sensitive element sequence; and calculate the replication probability according to the seed sequence and the corresponding sensitive element sequence, determine the conditional probability of the sensitive element sequence according to the generation probability and the replication probability, and complete the extraction of the sensitive element with the maximization of the conditional probability of the sensitive element sequence as the objective function.
[0099] A corpus module, configured to determine the relative offset of the sensitive element with respect to the data packet header of the seed sequence where it is located, determine the ID according to the sensitive element and the relative offset, determine the weight of the sensitive element according to the generation probability, code coverage rate, and status coverage rate of the sensitive element, and construct a corpus with the ID as the index, the sensitive element and its relative offset, weight, and response status code.
[0100] A testing module, configured to, for the protocol data packet to be mutated, select the corresponding corpus in the corpus according to the protocol state, and perform replacement mutation with the sensitive element and relative offset corresponding to the maximum weight, thereby obtaining a test case, and complete the fuzz testing process by monitoring the execution situation of the test case.
[0101] It should be noted here that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules can be executed in a computer system such as a set of computer-executable instructions as part of the system.
[0102] In more embodiments, there is also provided:
[0103] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0104] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0105] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0106] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0107] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0108] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.
[0109] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules can be located in local and remote storage media.
[0110] The computer program code for implementing the method of the present invention can be written in one or more programming languages. Such computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0111] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0113] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A grey-box fuzz testing method for SOME / IP protocol, characterized in that including: Calculating the generation probability for each element of each seed sequence in the obtained protocol data packet, taking the element with the maximum generation probability as the sensitive element, and thus determining the sensitive element sequence; Calculating the replication probability based on the seed sequence and the corresponding sensitive element sequence, determining the conditional probability of the sensitive element sequence according to the generation probability and the replication probability, and taking maximizing the conditional probability of the sensitive element sequence as the objective function to complete the extraction of sensitive elements; Determining the relative offset of the sensitive element relative to the data packet header of the seed sequence where it is located, determining the ID according to the sensitive element and the relative offset, determining the weight of the sensitive element according to the generation probability, code coverage rate, and status coverage rate of the sensitive element, and constructing a corpus with the ID as the index, the sensitive element and its relative offset, weight, and response status code; For the protocol data packet to be mutated, selecting the corresponding corpus in the corpus according to the protocol state, and performing replacement mutation with the sensitive element and relative offset corresponding to the maximum weight, thereby obtaining a test case, and completing the fuzz testing process by monitoring the execution of the test case.
2. The some / ip protocol grey-box fuzz testing method according to claim 1, wherein Generated probability p c is as follows: p c (SS t |SS <t ,S) = o(SS t-1 ,h t ,S′); h t = f(SS t-1 ,h t-1 ,S′); Among them, h t is the output of the hidden layer of the GRU network at time t, h t-1 is the output of the hidden layer of the GRU network at time t - 1, and f(·) is the calculation function of the hidden layer of the GRU network; SS t is all the sensitive element sequences extracted at time t, and SS t-1 is all the sensitive element sequences extracted at time t - 1. o(SS t-1 , h t , S′) is the generation probability calculation function. Specifically, at time t, it calculates the generation probability that each element of each seed sequence in the extraction vector S′ is a sensitive element, and takes the element with the maximum generation probability as the sensitive element at time t.
3. The some / ip protocol gray-box fuzz testing method according to claim 1, characterized in that The extraction vector S′ is: after obtaining the protocol data packet, encoding each seed sequence in the protocol data packet to obtain an extraction vector, and extracting sensitive elements based on the extraction vector.
4. The some / ip protocol gray-box fuzz testing method according to claim 1, characterized in that Replication probability p s is as follows: Among them, θ(SS t = S i ) is the output score calculated after the seed sequence S i is copied to SS t , that is, to determine the element positions of the seed sequence S i corresponding to SS t that belong to sensitive elements after that, and thus determine the output score through the function θ(SS t = S i ); I is the number of sequences in SS t ; Z is the sum of all scores; hh i is the operation of the output parameter matrix after the representation layer encoding of S i ; ω s is the weight parameter matrix; σ(·) is the tanh non-linear activation function; T is the matrix transpose.
5. A some / ip protocol grey-box fuzz testing method according to any one of claims 2 or 4, characterized in that Defining the objective function of maximizing the conditional probability of the sensitive element sequence as: P(SS t |SS <t ,S) = p c (SS t |SS <t ,S) + γp s (SS t |SS <t ,S); Among them, SS <t is all sensitive element sequences extracted before time t; γ is a harmonic parameter.
6. The some / ip protocol grey-box fuzz testing method according to claim 1, characterized in that The weight is: W(ID) = O(·) + Ψ(·) + Φ(·); where W is the weight; O(·) is the generation probability of the sensitive element under the current ID index, Ψ(·) is the code coverage rate of the sensitive element under the current ID index, Φ(·) is the status coverage rate of the sensitive element under the current ID index, and is obtained by dividing the number of relevant states of the sensitive element under the current ID index by the total number of states.
7. A SOME / IP protocol grey-box fuzz testing system, characterized in that, including: A sensitive structure extraction module configured to calculate the generation probability for each element of each seed sequence in the obtained protocol data packet, taking the element with the maximum generation probability as the sensitive element, and thus determining the sensitive element sequence; Calculating the replication probability based on the seed sequence and the corresponding sensitive element sequence, determining the conditional probability of the sensitive element sequence according to the generation probability and the replication probability, and taking maximizing the conditional probability of the sensitive element sequence as the objective function to complete the extraction of sensitive elements; A corpus module configured to determine the relative offset of the sensitive element relative to the data packet header of the seed sequence where it is located, determine the ID according to the sensitive element and the relative offset, determine the weight of the sensitive element according to the generation probability, code coverage rate, and status coverage rate of the sensitive element, and construct a corpus with the ID as the index, the sensitive element and its relative offset, weight, and response status code; A testing module configured to, for the protocol data packet to be mutated, select the corresponding corpus in the corpus according to the protocol state, and perform replacement mutation with the sensitive element and relative offset corresponding to the maximum weight, thereby obtaining a test case, and completing the fuzz testing process by monitoring the execution of the test case.
8. An electronic device, characterized in that, including a memory, a processor, and computer instructions stored on the memory and running on the processor, and when the computer instructions are run by the processor, implementing the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by a processor, the method according to any one of claims 1-6 is completed.
10. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Vulnerability type guiding fuzzy testing method and system based on byte sensitive energy distribution
CN114756471A
Network protocol fuzz testing method and system based on state and message awareness
CN116170181A
Network protocol grey box fuzzy test method and device based on equipment response state
CN119814353A