Protocol Vulnerability Mining Test Method and System Guided by Fine-Grained Coverage Information

Through the fine-grained coverage information guidance method, seeds with higher coverage frequency are preferred for mutation, which solves the problem of insufficient exploration capabilities of protocol entity program branches in the existing technology, and achieves more efficient protocol vulnerability discovery.

CN116827836BActive Publication Date: 2025-06-17HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310575664.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2025-06-17
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

The existing protocol fuzz testing method based on syntax generation and coverage information guidance has a coarse granularity in the utilization of coverage information, and has failed to effectively explore branches with low coverage frequency in protocol entity programs, resulting in insufficient vulnerability discovery capabilities.

Method used

The protocol vulnerability mining test method based on fine-grained coverage information is adopted. The weight of the seed is calculated by calculating the coverage frequency of the seed coverage branch path, and the seed with the highest weight is preferred for mutation. The test cases are guided to cover more rare branches and dynamically adjust the energy value of the seeds to eliminate seeds with poor performance.

Benefits of technology

The ability of fuzz testing to explore protocol entity program branches has been improved, more potential protocol vulnerabilities have been discovered, and the testing efficiency has been improved by quickly discarding seeds with poor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116827836B_ABST
    Figure CN116827836B_ABST
Patent Text Reader

Abstract

The present invention discloses a protocol vulnerability mining and testing method guided by fine-grained coverage information. During the testing process, the weight of the seed message is calculated according to the coverage frequency of the seed covering the branch path, and the seed message with the highest weight is preferentially selected for mutation to guide the newly generated test cases to cover more program branches; an initial energy value is assigned to each selected seed, and the energy value is dynamically adjusted according to the coverage information feedback, and then the seeds with low energy values are eliminated; finally, during the process of mutating the seeds, on the basis of the original mutation of the seeds, other related fields are continuously mutated to enable the generated test cases to explore and execute more program branches, improve the protocol test coverage rate, and discover more protocol vulnerabilities. The present invention can improve the test branch coverage rate of the protocol program and discover more protocol vulnerabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of protocol automated security testing, and particularly relates to a protocol vulnerability mining test method and system guided by fine-grained coverage information. Background Art

[0002] Protocol fuzz testing is an automated protocol vulnerability mining test technology. It inputs a large number of random, illegal or unexpected data to the protocol entity program to be tested, in order to discover its potential vulnerabilities or anomalies. In order to make the test cases generated during the network protocol fuzz testing process more compliant with the requirements of the protocol specification, improve the acceptance rate of the test cases, and at the same time improve the code coverage rate of the protocol entity program. In recent years, related research work has proposed fuzz testing methods based on syntax generation and coverage information guidance, which have solved the problem that a large number of message mutation operations during the protocol testing process destroy the structural format of the message, and the problem that the black-box protocol fuzzer based on syntax generation lacks coverage information feedback guidance.

[0003] Related work on protocol vulnerability mining testing based on syntax generation and coverage information guidance includes: Peach*, PAVFuzz, Z-Fuzzer, and EPF, etc. Z-Fuzzer introduces coverage information feedback on the black-box protocol fuzzer BooFuzz based on syntax generation to guide subsequent message mutation. EPF uses swarm-based simulated annealing to heuristically schedule the test cases in the seed library during the fuzzing process, recombines and mutates the test cases in the seed library to generate new test cases. PAVFuzz calculates and updates the mutation weights of each variable field by learning the relationship between two fields of different data models, guiding the fuzz testing towards the direction of maximizing the coverage rate. Peach* introduces a coverage feedback mechanism on the basis of Peach, and constructs higher-quality test cases by using the test cases that trigger the coverage of new paths.

[0004] However, the above-mentioned protocol fuzz testing methods based on syntax generation and coverage information guidance have a relatively coarse utilization granularity of the coverage information. They only consider the situation of coverage rate growth, and do not consider which branches of the protocol entity program are covered by the test cases. Some of these branches may be covered multiple times, and some branches may be covered very few times. If these information can be fully utilized to guide the fuzzer to explore more branches with lower coverage frequency during the fuzz testing process, it will improve the branch exploration ability of the fuzzer for the protocol entity program, and thus discover more potential protocol vulnerabilities. Summary of the Invention

[0005] The present invention proposes a protocol vulnerability mining test method based on fine-grained coverage information guidance. During the fuzzy test process, the weight of the seed is calculated according to the coverage frequency of the seed coverage branch path, and the seed with the highest weight is preferentially selected for mutation, guiding the newly generated test cases to cover more rare branches; an initial energy value is assigned to each selected seed, and the energy value is dynamically adjusted according to the coverage information feedback, and then the seeds with low energy values ​​are eliminated; finally, in the process of seed mutation, the original mutation of the seed is maintained, and other fields are mutated successively, so that the generated test cases can explore updated branches on the basis of triggering the original branches.

[0006] The method is divided into a preprocessing phase, a test preparation phase, and a fuzz testing phase. In the preprocessing phase, the entity program of the protocol to be tested is instrumented and compiled to obtain the corresponding binary executable program, and the session model of the protocol to be tested is defined as the input of the fuzz testing phase. In the test preparation phase, the execution engine is built, the global data structure required in the fuzz testing process is created, and the initial seed library corresponding to each data model is created. The fuzz testing phase includes processes such as seed selection, test case generation, and test case evaluation. Seed selection: The weight of each seed is calculated based on the fine-grained branch coverage information of each seed in the seed library, and the seeds in the seed library are sorted from large to small according to the weight, and the seed with the largest weight is preferentially selected as the benchmark for the current generated test case. Test case generation: Generate new test cases without destroying the beneficial mutation of the seeds, that is, keep the original conditions for triggering new branches of the seeds unchanged, and continue to mutate other variable fields. Test case evaluation: Check whether the test case triggers a crash or a new branch path. If a new branch path is triggered, the test case is stored in the corresponding seed library as a new seed together with its branch coverage information. If a crash is triggered, the test case is saved to facilitate the reproduction and analysis of the crash.

[0007] The present invention proposes a protocol vulnerability mining test method based on fine-grained coverage information guidance, which includes three main stages: a preprocessing stage, a test preparation stage, and a fuzzy testing stage; the fuzzy testing stage includes processes such as seed selection, test case generation, and test case evaluation.

[0008] 1. Preprocessing stage

[0009] 1-1. Instrumentation compilation of the protocol entity program source code

[0010] In the fuzz testing process, in order to obtain the branch coverage of the protocol entity program to be tested, it is necessary to insert and compile the protocol entity program source code using the gcc compilation tool that comes with aflfast to generate the corresponding binary executable program;

[0011] 1-2. Define a data model collection

[0012] By analyzing the request data packets corresponding to the protocol to be tested and combining with the protocol specification to be tested, use the data model definition function provided by the BooFuzz fuzz testing framework to define the set of protocol data models model set ={model1,..., model i ,..., model n}, i = 1,..., n, where n is the total number of protocol data models to be tested, and the defined data model is used as the protocol specification template for generating corresponding test cases;

[0013] 1-3. Define the session model

[0014] Combined with the requirements of the session message sequence in the protocol specification to be tested, use the session model definition function provided by the BooFuzz fuzz testing framework to connect the data models in the set of data models model set into the session model sessionModels = {sessionSeq1,..., sessionSeq j ,…, sessionSeq m}, j = 1,..., m, where sessionSeq j represents a session sequence in the session model sessionModels, and the session sequence is composed of several data models in the set of data models model set in a certain front-to-back order relationship, and m is the number of session sequences in the session model.

[0015] 2. Test preparation stage

[0016] 2-1. Build an execution engine

[0017] The execution engine is an independent module used to receive the test cases generated during the fuzz testing process, and is used to execute the binary executable program generated after instrumented compilation.

[0018] In order to record the coverage of each branch of the protocol entity program during the fuzz testing process, a shared memory shareMem with a size of XM bytes is allocated when building the execution engine (the size of shareMem is an empirical value, generally set to 64K bytes);

[0019] In order to separately record the branch coverage of the current test case, a temporary shared memory temShareMem with the same size of XM bytes is allocated. During the fuzz testing process, if the current test case can be used as a new seed, the information stored in the temporary shared memory can be used as the branch coverage information of the seed for the protocol entity program.

[0020] 2-2. Create the global data structures required during the fuzz testing process

[0021] In order to be able to count the number of times each branch is covered, create a global integer array branCovArr. Each array element counts the number of times a branch is covered. To correspond to the global shared memory, the length of the integer array is taken as XM. Each element is an integer, which is 4 bytes, and the maximum value is 2 32 -1, which is sufficient to count the number of times the corresponding branch is covered during the fuzz testing process. The memory size occupied by the global integer array branCovArr is 4 * XM bytes.

[0022] 2-3. System warm-up, initialize the data model seed bank

[0023] 2-3-1. Traverse the session models sessionModels defined in step 1-3, and sequentially traverse and take out the session sequences sessionSeq j , and turn to step 2-3-2; if the traversal of the session models has ended, the system warm-up work has been completed.

[0024] 2-3-2. According to the order relationship of the data models in the session sequence sessionSeq j obtained in step 2-3-1, traverse the session sequence sessionSeq j to obtain the data model model i , and turn to step 2-3-3; if the traversal work for the session sequence sessionSeq j has been completed, reset the protocol entity program to the initial state and turn to step 2-3-1.

[0025] 2-3-3. Generate an unmutated test case according to the data model model i obtained by traversing in step 2-3-2, and use it as the initial seed seed i in the seed bank seedBank i of the data model model i 0 (representing the 0th seed in the seed bank seedBank i ), input it into the execution engine, and create an array seedBranCovArr of branch coverage information corresponding to this seed i 0 , traverse the temporary shared memory temShareMem. If the corresponding subscript position y in the temporary shared memory is not zero, that is, the branch represented by the subscript at the corresponding position is covered, add the branch represented by this subscript y to the array seedBranCovArr of branch coverage information of this seed i0 In the process of traversing the temporary shared memory temShareMem, the global integer array branCovArr is updated. If the value at the corresponding subscript position y in the temporary shared memory is not zero, the data value stored in branCovArr[y] is incremented by one to count and update the number of times each branch is covered. Finally, the seed i 0 and its corresponding branch coverage information array seedBranCovArr i 0 are added to the data model model i corresponding to the seed bank seedBank i In it, seedBanki = {[seedi0, seedBranCovArr i 0}, and the temporary shared memory temShareMem is cleared. Then turn to step 2-3-2 to initialize the seed bank for the next data model.

[0026] 3. Fuzz testing phase

[0027] 3-1. Data model selection

[0028] 3-1-1. Traverse the session models sessionModels defined in step 1-3, and sequentially traverse and take out the session sequence sessionSeq j , and then turn to step 3-1-2; if the traversal of the session models has ended, then end the fuzz testing work.

[0029] 3-1-2. According to the order relationship of the data models in the session sequence sessionSeq obtained in step 3-1-1 j , traverse the session sequence sessionSeq j to obtain the data model model i , which is used as the data model in the current fuzz testing phase, and then turn to step 3-2; if the traversal of the session sequence sessionSeq j has been completed, then turn to step 3-1-1 to select a new session sequence.

[0030] 3-2. Seed selection

[0031] Using the seed bank seedBank corresponding to the data model model selected in step 3-1-2 i i ​The fine-grained branch coverage information array seedBranCovArr for each seed in calculates the weight of each seed, and sorts the seeds in the seed bank in descending order of weight, and preferentially selects the seed with the largest weight for the current fuzz testing phase. The specific steps are as follows:

[0032] 3-2-1. Traverse the seed bank seedBank i and calculate the seed weight

[0033] If there is no seed in the seed bank seedBank i it means that the data model model i has been fully fuzz tested at this stage, and turn to step 3-1-2 to select a new data model in the current session sequence sessionSeq j as the data model used at this stage;

[0034] If there is only one seed in the seed bank seedBank i there is no need to sort the seeds, and directly turn to step 3-2-3 to select this seed as the benchmark for generating new test cases at this stage;

[0035] Otherwise, create the weight array seedWeight i corresponding to the seed bank seedBank i = [], with an initial value of empty. Traverse the seed bank seedBank corresponding to the data model model selected in step 3-1-2 i = {[seed i ,seedBranCovArr i 0 ,...,[seed i 0 ,...,[seed i k ,seedBranCovArr i k ,...,[seed i t ,seedBranCovArr i t},k≤1,...,t, where t is the number of seeds in the seed bank seedBank i Select a seed seed i k in turn from it, and according to the branch coverage information array seedBranCovArr of the seed seed i k calculate the seed seed i k weighti k The weight of i k , the seed branch coverage information array seedBranCovArr i k records the seed i k which branches of the protocol entity program are covered. The formula for calculating the weight of the seed is as follows:

[0036]

[0037] where weight i k represents the weight of the seed i k , seedBranCovArr i k is the branch coverage information array of this seed, recording which branches this seed covers. The length of this array len(seedBranCovArr i k ) is the number of branches covered by this seed. Let r range from 0 to length(seedBranCovArr i k ) - 1 to traverse the array seedBranCovArr i k , where branCovArr[seedBranCovArr i k [r]] is the number of times the branch represented by seedBranCovArr i k [r] is covered during the fuzz testing process. Let it be w, and 1 / w 2 will rapidly approach 0 as w increases, that is, the lower the frequency of the branch covered by the seed during the fuzz testing process, the greater the contribution of this branch to the weight i k value of the seed. After calculating the weight i k of the seed i k , store its weight information into the weight array seedWeight i corresponding to the seed bank seedBank i = [..., weight i k ,...].

[0038] Traverse the seed bank seedBank in the same way i Save the weight of each seed calculated during the process into the seed bank seedBank i The corresponding weight array seedWeight i After the traversal of the seed bank is completed, the seed weight array seedWeight is obtained i = [weight i 0 ,..., weight i k ,..., weight i t , k = 1,..., t, where t is the number of seeds in the seed bank seedBank i Go to step 3-2-2

[0039] 3-2-2. Sort the seeds in the seed bank seedBank i Sort the seeds

[0040] Using the seed bank seedBank obtained in step 3-2-1 i The corresponding weight array seedWeight i = [weight i 0 ,..., weight i k ,..., weight i t , sort the corresponding seeds in the seed bank seedBank i = {[seed i 0 , seedBranCovArr i 0 ,..., [seed i k , seedBranCovArr i k ,..., [seed i t , seedBranCovArr i t} from largest to smallest by weight

[0041] 3-2-3. Select seeds

[0042] From the seed bank seedBank sorted in step 3-2-2 iSelect the first seed, that is, the seed with the largest weight, as the benchmark for generating test cases in the current stage. Assume that the seed with the largest weight is seed i k . And initialize the energy value of the seed seedEnergy i k =SC, SC value is 1, initialize the threshold of the seed i k =0.0001. Go to step 3-3 and use the seed i k Execute the test case generation process.

[0043] 3-3. Test case generation

[0044] Without destroying the optimal seed selected in step 3-2 i k Generate new test cases based on beneficial mutations, that is, keep the original seed triggering new branches unchanged, and continue to mutate other variable fields on this basis to improve the pertinence of the newly generated test cases to explore deeper code levels. The specific steps are as follows:

[0045] 3-3-1. Get the current seed i k Corresponding data model model i All mutable fields in the seed i k The value in fields = {field1, ..., field p , ..., field q}, p = 1, ..., q, where q is the data model model i The total number of mutable fields in the . And the initialization mutated field position mark bit index b =0.

[0046] 3-3-2. Traverse the values ​​of all variable fields in the seed obtained in step 3-3-1, assuming that the current traversed field value is field p If you have reached the last field of the seed while traversing fields, reset the index. b If it is 0, restart the traversal and make the next round of mutation start from the first unmutated field.

[0047] 3-3-3. If the field value traversed in step 3-3-2 is field p It's a seed i kThe value that has not undergone mutation, and this field is at the marker bit index b For the fields after that, use the mutation engine of EPF to mutate this field rield p to obtain the current field field p The mutation result mutated_field of p , and set the marker bit index b to p; otherwise, go to step 3-3-2 to continue traversing fields.

[0048] 3-3-4. Use the mutation result mutated_field of the current field field obtained in step 3-3-3 p to replace the corresponding field value of the seed seed p to obtain the newly generated test case testcase. i k

[0049] 3-3-5. Based on the current session sequence sessionSeq j obtain all the pre-data models of the data model model i , then generate the unmutated pre-message sequence preMessSeq according to the pre-data models, combine the test case testcase generated in step 3-3-4 with the pre-message sequence preMessSeq to form the message sequence messSeq, and inject the message sequence messSeq into the execution engine. Go to step 3-4 for test case evaluation.

[0050] 3-4. Test case evaluation

[0051] Monitor the status of the execution engine and evaluate the current test case. If the message sequence messSeq containing the new test case causes the execution engine to crash, then generate a crash report and save the test case testcase that triggered the crash; and by analyzing the global shared memory, evaluate whether this test case triggers a new branch coverage. If a new branch coverage is triggered, use this test case as a seed, the branch coverage information of this test case is stored in the temporary shared memory, traverse the temporary shared memory, count the branch coverage information of this test case and store it in the seed library together with this test case. The specific steps are as follows:

[0052] 3-4-1. Monitor the status of the execution engine

[0053] Monitor the status of the execution engine. If the message sequence messSeq containing the new test case causes the execution engine to crash, then generate a crash report and save the test case testcase that triggered the crash for post-mortem analysis after fuzz testing.

[0054] 3 - 4 - 2. Evaluate whether a new branch is triggered

[0055] Traverse the global shared memory shareMem to determine whether the currently executed test case testcase has triggered a new branch. If there are new non - zero bytes in the global shared memory shareMem, it means a new branch has been triggered.

[0056] If a new branch is triggered, use the current test case as a new seed i x , and create an array of branch coverage information seedBranCovArr corresponding to this seed i x , traverse the temporary shared memory temShareMem. If the corresponding subscript position y in the temporary shared memory is not zero, that is, the branch represented by the subscript at the corresponding position is covered, add the branch represented by this subscript to the array of branch coverage information seedBranCovArr of this seed i x During the process of traversing the temporary shared memory temShareMem, update the global integer array branCovArr. If the corresponding subscript position y in the temporary shared memory is not zero, then increment the data value stored in branCovArr[y] by one to count and update the number of times each branch is covered. Finally, add the seed i x and its corresponding array of branch coverage information seedBranCovArr i x to the current data model model i corresponding seed bank seedBank i and clear the temporary shared memory temShareMem.

[0057] 3 - 4 - 3. Update and check the seed energy value

[0058] According to the evaluation result in step 3 - 4 - 2, adjust the energy value seedEnergy of the current stage template seed seed selected in step 3 - 2 - 3 i k of the seed i k . If the evaluation result is that a new branch is triggered, update the energy value seedEnergy of the seed i k to the initial value SC; otherwise, update the energy value seedEnergy of the seed i k to i k the energy value seedEnergy of the seedi k Updated to seedEnergy i k ×α (α is the attenuation factor, the value is less than 1 and greater than 0. The smaller the value, the greater the attenuation of the seed energy value. The default value of this method is 0.97);

[0059] After the update, if the current template seed is i k Energy value of seedEnergy i k <threshold i k , it represents the seed i k The performance of the seed library is poor. i Delete the seed and the branch coverage information array of the seed, and go to step 3-2 to remove the seed from the seed library seedBank i Select a new seed as the template seed for the current stage; if the seed energy value seedEnergy i k >=threshold i k , then go to step 3-3-2 and continue to use the seed to generate new test cases.

[0060] Another object of the present invention is to provide a protocol vulnerability mining test system FCIFuzz (Fine-grained Coverage Information Fuzzer) based on fine-grained coverage information guidance, such as Figure 1 As shown, including:

[0061] Preprocessing module. The main function of the preprocessing module is to use the gcc compiler tool provided by aflfast to perform stub compilation on the entity program of the protocol to be tested, and obtain the corresponding binary executable program; use the data model definition function provided by the BooFuzz ​​fuzz testing framework to define the data model set model of the protocol set , combined with the requirements of the protocol specification to be tested for the session message sequence, the defined data model set model set The data models in are connected into session models sessionModels as input for the fuzz testing phase.

[0062] Test preparation module. The main functions of the test preparation module are to build the execution engine, create the global integer array branCovArr required in the fuzz testing process, and create the initial seed library corresponding to each data model.

[0063] Fuzz testing module. The fuzz testing module sequentially selects session sequences from the session models defined by the preprocessing module, and then sequentially selects data models in the session sequences to obtain the data model model for the current stage of testing. i , and finally execute the seed selection module, test case generation module, and test case evaluation module among them:

[0064] Seed selection module. This module calculates the weight weight of the seed seed according to the branch coverage information array seedBranCovArr of the seed corresponding to the data model model i in the corresponding seed bank seedBank i and the global integer array branCovArr created in the test preparation stage, i k and sorts the seeds in the seed bank in descending order according to the weight, preferentially selects the seed with the largest weight as the benchmark for generating test cases in the current stage, and initializes the energy value seedEnergy of the selected seed i k = SC. i k i k i k

[0065] Test case generation module. This module generates new test cases testcase on the basis of not destroying the original beneficial mutations of the seed seed i k selected by the seed selection module. And form a message sequence messSeq with the generated test cases and the corresponding prefix message sequence preMessSeq, and inject the message sequence messSeq into the execution engine. i k

[0066] Test case evaluation module. This module monitors the reaction of the execution engine. If the message sequence messSeq containing the new test case testcase causes the execution engine to crash, a crash report is generated and the test case testcase that triggers the crash is saved. And by analyzing the global shared memory shareMem, it evaluates whether the test case testcase triggers a new branch coverage. If a new branch coverage is triggered, the test case is used as a new seed. The branch coverage information of the new seed is stored in the temporary shared memory temShareMem. Traverse the temporary shared memory, count the branch coverage information of the new seed to obtain the branch coverage information array seedBranCovArr corresponding to the new seed, and store it together with the new seed into the current data model model i in the seed bank seedBank i ; And according to the evaluation result, update and check the current template seed seed i k energy value seedEnergy i k .

[0067] The beneficial effects of the present invention are as follows:

[0068] 1. The present invention adopts a protocol fuzzing testing method based on grammar generation and fine-grained coverage information guidance, considering the coverage of each seed for branches with a lower coverage frequency, guiding the fuzzing testing work to cover more branches with a lower coverage frequency, and improving the ability to fully discover vulnerabilities during the fuzzing testing process.

[0069] 2. In order to be able to discard seeds with poor performance faster, the present invention gives each seed an initial energy value. During the fuzzing testing process, the energy value is dynamically adjusted to achieve the purpose of faster discarding of seeds with poor performance. And during the mutation process of seeds with better performance, the original beneficial mutations of the seeds are kept unchanged, and other fields are mutated one after another, so that the newly generated test cases can explore newer branches while still triggering the original branches as much as possible.

[0070] 3. Compared with the existing protocol fuzzing testing tools based on grammar generation and coverage information guidance, the protocol fuzzing testing method based on fine-grained coverage information guidance proposed by the present invention can effectively solve the problem that the existing protocol fuzzing testing tools based on grammar generation and coverage information guidance have a coarser utilization granularity of the coverage information of test cases and cannot more effectively guide the fuzzing testing to cover more valuable program branches. It improves the exploration ability and vulnerability discovery ability for the target program branches during the fuzzing testing process. Description of the Drawings

[0071] Figure 1, a flowchart of the protocol vulnerability mining test method guided by fine-grained coverage information;

[0072] Figure 2 , Taking RTSP protocol as an example, the diagram of the defined session model;

[0073] Figure 3 , Schematic diagram of seed weight calculation during seed selection;

[0074] Figure 4 , executing the test case generation module multiple times with the same seed, and generating a test case diagram;

[0075] Figure 5 , the trend chart of the number of branch coverage changes on the target protocol entity program. DETAILED DESCRIPTION

[0076] The technical solution of the present invention will be fully described below in conjunction with the accompanying drawings.

[0077] like Figure 1 As shown, the overall steps of the protocol vulnerability mining test method guided by fine-grained coverage information are divided into three main stages: preprocessing stage, test preparation stage, and fuzz testing stage; the fuzz testing stage includes processes such as seed selection, test case generation, and test case evaluation.

[0078] 1. The preprocessing stage includes the following steps:

[0079] 1-1. Instrumentation compilation of the protocol entity program source code

[0080] In the fuzz testing process, in order to obtain the branch coverage of the protocol entity program to be tested, it is necessary to insert and compile the protocol entity program source code using the gcc compilation tool that comes with aflfast to generate the corresponding binary executable program;

[0081] 1-2. Define a data model collection

[0082] By analyzing the request data packets corresponding to the protocol to be tested and combining the protocol specification to be tested, the data model definition function provided by the BooFuzz ​​fuzz testing framework is used to define the data model set model of the protocol set ={model1, ..., model i , ..., model n}, i=1, ..., n, where n is the total number of protocol data models to be tested, and the defined data model is used as a protocol specification template for generating corresponding test cases;

[0083] 1-3. Define the conversation model

[0084] Combined with the requirements of the session message sequence for the protocol to be tested, use the session model definition function provided by the BooFuzz fuzz testing framework to connect the data model set model defined in steps 1-2 set in the data models in the session model sessionModels = {sessionSeq1,..., sessionSeq j , …, sessionSeq m}, j = 1,..., m, where sessionSeq j represents a session sequence in the session model sessionModels. The session sequence is composed of several data models in the data model set model set in a certain order before and after, and m is the number of session sequences in the session model.

[0085] Taking the RTSP protocol as an example, the schematic diagram of the defined session model structure is as shown in Figure 2 the figure. This session model has 5 session sequences, and each element in the session sequence is an independent data model.

[0086] 2. The test preparation stage includes the following steps:

[0087] 2-1. Build an execution engine

[0088] The execution engine is an independent module used to receive the test cases generated during the fuzz testing process to execute the binary executable program generated after instrumented compilation.

[0089] In order to record the coverage of each branch of the protocol entity program during the fuzz testing process, a shared memory shareMem with a size of XM bytes is allocated when building the execution engine (the size of shareMem is an empirical value, generally set to 64K bytes);

[0090] In order to separately record the branch coverage of the current test case, a temporary shared memory temShareMem with the same size of XM bytes is allocated. During the fuzz testing process, if the current test case can be used as a new seed, the information stored in the temporary shared memory can be used as the branch coverage information of the seed for the protocol entity program.

[0091] 2-2. Create the global data structures required during the fuzz testing process

[0092] In order to be able to count the number of times each branch is covered, a global integer array branCovArr is created. Each array element counts the number of times a branch is covered. To correspond to the global shared memory, the length of the integer array is taken as XM, each element is an integer, 4 bytes, and the maximum value is 2 32-1, which is sufficient to count the number of times the corresponding branch is covered during the fuzz testing process. The memory size occupied by the global integer array branCovArr is 4 * XM bytes.

[0093] 2 - 3. System warm-up, initialize the data model seed bank

[0094] 2 - 3 - 1. Traverse the session models sessionModels defined in steps 1 - 3, sequentially traverse and extract the session sequences sessionSeqj, and turn to step 2 - 3 - 2; if the traversal of the session models is completed, the system warm-up work has been completed.

[0095] 2 - 3 - 2. According to the order relationship of the data models in the session sequence sessionSeq obtained in step 2 - 3 - 1 j in the session sequence sessionSeq j obtain the data model model i , and turn to step 2 - 3 - 3; if the traversal of the session sequence sessionSeq j is completed, reset the protocol entity program to the initial state and turn to step 2 - 3 - 1.

[0096] 2 - 3 - 3. According to the data model model obtained by traversing in step 2 - 3 - 2 i generate an unmutated test case, and use it as the initial seed seed i in the seed bank seedBank i of the data model model i 0 (representing the 0th seed in the seed bank seedBank i ), input it into the execution engine, and create an array seedBranCovArr for branch coverage information corresponding to this seed i 0 , traverse the temporary shared memory temShareMem. If the value at the corresponding subscript position y in the temporary shared memory is not zero, that is, the branch represented by the subscript at the corresponding position is covered, add the branch represented by the subscript y to the branch coverage information array seedBranCovArr i 0 of this seed, and update the global integer array branCovArr during the traversal of the temporary shared memory temShareMem. If the value at the corresponding subscript position y in the temporary shared memory is not zero, then the data value stored in branCovArr[y] is incremented by one to count and update the number of times each branch is covered. Finally, the seed seed i 0and its corresponding array of branch coverage information for seeds, seedBranCovArr i 0 Add it to the data model, model i The corresponding seed bank, seedBank i In seedBank i = {[seed i 0 , seedBranCovArr i 0}, and clear the temporary shared memory, temShareMem. Go to step 2-3-2 to initialize the seed bank for the next data model.

[0097] 3. The fuzz testing phase includes the following steps:

[0098] 3-1. Data model selection

[0099] 3-1-1. Traverse the session models, sessionModels, defined in step 1-3, and sequentially traverse and extract the session sequences, sessionSeq j , and go to step 3-1-2; if the traversal of the session models has ended, then end the fuzz testing work.

[0100] 3-1-2. According to the order relationship of the data models in the session sequence, sessionSeq, obtained in step 3-1-1, traverse the session sequence, sessionSeq j to obtain the data model, model j as the data model used in the current fuzz testing phase, and go to step 3-2; if the traversal of the session sequence, sessionSeq i has been completed, then go to step 3-1-1 to select a new session sequence. j

[0101] 3-2. Seed selection

[0102] Use the array of fine-grained branch coverage information for each seed, seedBranCovArr, in the seed bank, seedBank, corresponding to the data model selected in step 3-1-2 i to calculate the weight of each seed, and sort the seeds in the seed bank in descending order of weight, and preferentially select the seed with the largest weight for use in the current fuzz testing phase. The specific steps are as follows: i

[0103] 3-2-1. Traverse the seed bank, seedBank i to calculate the seed weights

[0104] If the seed bank i If there is no seed in the data model i At this stage, we have already undergone sufficient fuzz testing, and now we turn to step 3-1-2. In the current session sequence sessionSeq j Select the new data model as the data model used at this stage;

[0105] If the seed bank i If there is only one seed in the test case, there is no need to sort the seeds, and the process can be turned directly to step 3-2-3 to select the seed as the benchmark for generating new test cases at this stage.

[0106] Otherwise, create a seed library seedBank i The corresponding weight array seedWeight i = [], the initial value is empty. Traverse the data model selected in step 3-1-2 i The corresponding seed bank seedBank i ={[seed i 0 ,seedBranCovArr i 0 ],...,[seed i k ,seedBranCovArr i k ],...,[seed i t ,seedBranCovArr i t ]}, k = 1, ..., t, where t is the seed bank seedBank i The number of seeds in the table, from which one seed is selected in turn i k , and according to the seed i k The branch coverage information array seedBranCovArr i k Calculate the seed i k weight i k , the seed branch coverage information array seedBranCovArr i k Records the seed i k Which branches of the protocol entity program are covered. The formula for calculating the weight of the seed is as follows:

[0107]

[0108] Where weight i k Represents seed i k The weight of seedBranCovArr i k This is the branch coverage information array of the seed, which records which branches the seed covers. The length of the array is len (seedBranCovArr i k ) is the number of branches covered by the seed, and r takes the value from 0 to length(seedBranCovArr i k )-1 traverse the array seedBranCovArr i k , where branCovArr[seedBranCovArr i k [r]] is seedBranCovArr i k [r] represents the number of times the branch is covered during the fuzz testing process, set it to w, 1 / w 2 As w increases, it quickly approaches 0, that is, the lower the frequency of the branch covered by the seed during the fuzz testing process, the greater the weight of the branch on the seed. i k The greater the contribution of the value. Calculate the seed i k weight i k After that, the weight information is stored in the seed library seedBank i The corresponding weight array seedWeight i =[…,weight i k , ...].

[0109] In the same way, we will traverse the seed library seedBank i The weight of each seed calculated in the process is stored in the seed bank seedBank i The corresponding weight array seedWeight i , when the traversal of the seed library is completed, the seed weight array seedWeight is obtained i =[weight i 0 , ..., weighti k ,..., weight i t , k = 1,..., t, where t is the number of seeds in the seed bank seedBank i , and go to step 3-2-2.

[0110] 3-2-2. Sort the seeds in the seed bank seedBank i in the seed bank

[0111] Using the seed bank seedBank obtained in step 3-2-1 i corresponding weight array seedWeight i = [weight i 0 , …, weight i k ,..., weight i t , sort the corresponding seeds in the seed bank seedBank i = {[seed i 0 , seedBranCovArr i 0 ,..., [seed i k , seedBranCovArr i k ,..., [seed i t , seedBranCovArr i t} in descending order of weight.

[0112] 3-2-3. Select a seed

[0113] Select the first seed from the seed bank seedBank sorted in step 3-2-2 i , that is, the seed with the largest weight as the benchmark for generating test cases in the current stage. Assume the selected seed with the largest weight is seed i k . And initialize the energy value seedEnergy of this seed i k = SC, where SC takes the value of 1, and initialize the threshold threshold of this seed i k = 0.0001. Go to step 3-3 and use this seed seed i kExecute the test case generation process. The pseudo-code of the algorithm for seed selection is as follows:

[0114]

[0115] For the seed bank seedBank i An example of calculating the seed weights is as Figure 3 shown. Assume that the current data model model i corresponds to a seed bank with five seeds numbered 1 - 5. Among them, the branch labels covered by seed 1 in the protocol entity program are [4, 6, 7, 8, 9], and the branch ratio covered by seed 2 is [5, 7, 4, 8, 9]. Assume that the global integer array branCovArr that records the number of times each branch of the protocol entity program is covered during the fuzz testing process is as Figure 3 shown in the lower part. The branches labeled 1 and 6 are rare branches and are only covered once. The calculated result of the weight weight1 of seed 1 using the formula for calculating the seed weight is 1.26233, and the calculated result of the weight weight2 of seed 2 is 0.51233, which proves that rare branches have a greater impact on the seed weight result. After calculating the weights of all seeds in the seed bank, the seed label result of sorting the seeds from largest to smallest by weight is [4, 1, 2, 3, 5], and the seed labeled 4 has the largest weight.

[0116] 3 - 3. Test case generation

[0117] Generate new test cases without destroying the beneficial mutations of the optimal seed seed i k selected in step 3 - 2, that is, keep the conditions for the seed to trigger new branches unchanged, and on this basis, continue to mutate other mutable fields to improve the pertinence of the newly generated test cases to explore deeper code levels. The specific steps are as follows:

[0118] 3 - 3 - 1. Obtain the current seed seed i k and all the mutable fields in the corresponding data model model i in the values fields = {field1,..., field i k ,..., field p ,..., field q} in the seed seed i , where p = 1,..., q, and q is the total number of mutable fields in the data model model b = 0. And initialize the position marker index

[0119] 3-3-2. Traverse all the values of the variable fields in the seed, fields, and assume that the currently traversed field value is field p . If, during the traversal of fields, the last field position in the seed has been reached, then reset the marker index b to 0 and start traversing again, so that the next round of mutation starts from the first unmutated field.

[0120] 3-3-3. If the field value field p traversed in step 3-3-2 i k is a value in the seed b that has not been mutated, and this field is a field after the marker index p , then use the mutation engine of EPF to mutate this field field p to obtain the mutated result mutated_field p of the current field, and set the marker index b to p; otherwise, go to step 3-3-2 to continue traversing fields.

[0121] 3-3-4. Use the mutated result mutated_field p of the current field obtained in step 3-3-3 p to replace the corresponding field value in the seed i k to obtain the newly generated test case testcase.

[0122] 3-3-5. According to the current session sequence sessionSeq j obtain all the pre-data models of the data model i , then generate the pre-message sequence preMessSeq that has not been mutated according to the pre-data models, form the message sequence messSeq by combining the test case testcase generated in step 3-3-4 with the pre-message sequence preMessSeq, and inject the message sequence messSeq into the execution engine. Go to step 3-4 for test case evaluation.

[0123] The pseudo-code of the algorithm flow for test case generation is as follows:

[0124]

[0125]

[0126] Use the seed ik After the test case generation module is executed multiple times, the results of the generated test cases are as follows Figure 4 shown. In the first round of mutation for this seed, all mutable fields in the seed are traversed. Fields that have already been mutated are skipped directly, and the next field is mutated to generate Figure 4 the test case testcase1 in; for the second round of mutation for this seed seed i k When performing the mutation, all mutable fields in the seed are traversed. Fields that have already been mutated are skipped directly, and the next field is checked. According to the mutation field position marker index b it can be seen that the next field has already been mutated, so continue to check the next field and mutate this field to generate the test case testcase2; the generation processes of test cases testcase3, testcase4, and testcase5 are the same; since the last mutable field has been mutated when generating test case testcase5, reset the mutation field position marker index b , in the next round of test case generation, a test case testcase6 that has mutated the same fields as test case testcase1 will be obtained, but the content of the mutation of the same fields is different for them.

[0127] 3-4. Test case evaluation

[0128] Monitor the status of the execution engine and evaluate the current test case. If the message sequence messSeq containing the new test case causes the execution engine to crash, a crash report is made, and the test case testcase that triggers the crash is saved; and by analyzing the global shared memory, it is evaluated whether this test case triggers a new branch coverage. If a new branch coverage is triggered, this test case is used as a seed, and the branch coverage information of this test case is stored in the temporary shared memory. Traverse the temporary shared memory, count the branch coverage information of this test case and store it in the seed library together with this test case. The specific steps are as follows:

[0129] 3-4-1. Monitor the status of the execution engine

[0130] Monitor the status of the execution engine. If the message sequence messSeq containing the new test case causes the execution engine to crash, a crash report is made, and the test case testcase that triggers the crash is saved for post-mortem analysis after the fuzz testing is completed.

[0131] 3-4-2. Evaluate whether a new branch is triggered

[0132] Traverse the global shared memory shareMem to determine whether the currently executed test case testcase has triggered a new branch. If there are new non-zero bytes in the global shared memory shareMem, it means that a new branch has been triggered.

[0133] If a new branch is triggered, use the current test case as a new seed i x , and create an array of branch coverage information seedBranCovArr corresponding to this seed i x , traverse the temporary shared memory temShareMem. If the corresponding subscript position y in the temporary shared memory is not zero, that is, the branch represented by the subscript at the corresponding position is covered, add the branch represented by this subscript to the array of branch coverage information seedBranCovArr of this seed i x . During the process of traversing the temporary shared memory temShareMem, update the global integer array branCovArr. If the corresponding subscript position y in the temporary shared memory is not zero, then increment the data value stored in branCovArr[y] by one to count and update the number of times each branch is covered. Finally, add the seed i x and its corresponding array of branch coverage information seedBranCovArr i x to the current data model model i corresponding to the seed bank seedBank i , and clear the temporary shared memory temShareMem.

[0134] 3-4-3. Update the check seed energy value

[0135] According to the evaluation result in step 3-4-2, adjust the energy value seedEnergy of the current stage template seed seed selected in step 3-2-3 i k of i k . If the evaluation result is that a new branch is triggered, update the energy value seedEnergy of the seed i k to the initial value SC; otherwise, update the energy value seedEnergy of the seed i k to i k seedEnergy i k updated to ik × a (where a is the attenuation factor, with a value less than 1 and greater than 0. The smaller the value, the greater the attenuation amplitude of the seed energy value. The default value of this method is 0.97);

[0136] After the update, if the energy value seedEnergy of the current template seed i k of seed i k < threshold i k , it means that the performance of this seed i k is not good. Delete this seed and the branch coverage information array of this seed from the seed bank i seedBank, and go to step 3-2 to select a new seed from the seed bank i seedBank as the current template seed; if the energy value seedEnergy of this seed i k >= threshold i k , then go to step 3-3-2 to continue generating new test cases using this seed.

[0137] The pseudo-code of the algorithm process for test case evaluation is as follows:

[0138]

[0139]

[0140] The above has described the embodiments of the present invention in detail with reference to the accompanying drawings. All equivalent changes and modifications made within the scope of the invention claimed in this application belong to the scope of protection of the present invention.

[0141] 4. Experimental verification

[0142] To prove the effectiveness of this method, the following conducts an experimental evaluation of the system FCIFuzz corresponding to this method in terms of three experimental indicators: the number of branch covers covered, the number of crashes triggered, and the number of seeds generated.

[0143] 4-1. Experimental design

[0144] Selection of experimental group and control group. The fuzzer FCIFuzz implemented based on this method is used as the experimental group, and the fuzzers EPF and PAVFuzz based on grammar generation and coverage information guidance are selected as the control groups. It should be noted that PAVFuzz is not open source, and this article reproduces it based on its method and uses it as the control group.

[0145] Table 1 Test Objectives

[0146] target program protocol description Live555 RTSP Real-Time Streaming Protocol Dnsmasq DNS Domain Name System

[0147] Test objective selection. As shown in Table 1, the RTSP protocol and DNS protocol are selected for experiments in this paper. Among them, the RTSP protocol is a real-time media streaming protocol. For this protocol, the Live555 protocol entity program is selected as the target program. The DNS protocol is a domain name resolution protocol. For this protocol, the Dnsmasq protocol entity program is selected as the target program. These protocol entity programs are open-source servers commonly used in reality and have practical evaluation value.

[0148] Experimental evaluation metrics. To evaluate the effectiveness of this method, the following three evaluation metrics are adopted in this paper:

[0149] (1) Number of branch covers. The number of branches of the protocol entity program that the fuzzer can cover within a specified time. This metric can reflect the exploration degree of the fuzzer on the code of the target protocol entity program. Only by covering more program branches is it possible to discover more potential vulnerabilities.

[0150] (2) Number of crashes triggered. The number of times the fuzzer causes the protocol entity program to crash within a specified time. This metric can directly reflect the vulnerability discovery ability of the fuzzer.

[0151] (3) Number of seeds generated. During the fuzz testing process, valuable test cases that trigger new branches are used as seeds. The more seeds there are, the more times the fuzzer triggers new branches, which can reflect the exploration ability of the fuzzer on the branches of the target program.

[0152] Experimental environment settings. The system environment of the experiment is the Ubuntu 20.04.4 LTS operating system, with 4GB of memory, a 2-core Intel(R) Core(TM) i5-9300HF CPU @ 2.40GHz 2.40GHz, and the experimental time for each group is 24 hours. To avoid the influence of randomness in the experiment, each group of experiments is repeated three times, and the subsequent chart data uses the average value of the results of the three experiments.

[0153] 4-2. Experimental Results

[0154] 4-2-1. Comparison of Branch Cover Number Results

[0155] Table 2 Results of Covered Branch Numbers

[0156]

[0157] Table 2 shows the results of the number of branch coverage of the experimental group FCIFuzz and the control groups EPF and PAVFuzz on the target programs. It can be seen from this table that the fuzzer FCIFuzz implemented based on this method covers more branches than the control groups EPF and PAVFuzz on the target programs Dnsmasq and Live555. Specifically, on the target program Dnsmasq, FCIFuzz covers 39.3 more program branches on average than EPF, with an improvement rate of 6.52%; and covers 52.3 more program branches on average than PAVFuzz, with an improvement rate of 8.86%. On the target program Live555, FCIFuzz covers 197.6 more program branches on average than EPF, with an improvement rate of 7.93%; and covers 274.6 more program branches on average than PAVFuzz, with an improvement rate of 11.37%. Finally, in terms of the average number of total branches covered by the target programs, FCIFuzz has an improvement of 7.66% compared to EPF and 10.88% compared to PAVFuzz. The improvement in the number of branch coverage of the target programs benefits from the fact that this method improves the seed scheduling algorithm of EPF, calculates the seed weights using fine-grained coverage information, and can preferentially use the seeds with the highest weights for mutation; and this method will keep the existing beneficial mutations of the seeds unchanged when generating new test cases, and explore new program branches on this basis.

[0158] During the testing process of the experimental group and the control groups on the target programs, the changing trend of the number of branch coverage is as Figure 5 shown. It can be seen from this figure that the experimental group FCIFuzz has a higher number of branch coverage of the target programs than the control groups EPF and PAVFuzz at the beginning because FCIFuzz has a warm-up process. It will first traverse the state model and generate a message sequence without mutation, and this process will cover most of the program branches. During the testing process of the control group EPF, there is an obvious step-growth phenomenon in the growth of the number of branch coverage. The reason for this phenomenon is that EPF uses a certain data model for sufficient testing before turning to the next data model, and the next data model may be unused, and the generated test messages of a new category will cover more branches in a short time, resulting in the step-growth phenomenon.

[0159] 4-2-2. Comparison of the results of the number of triggered crashes

[0160] Table 3 Results of the number of triggered crashes

[0161]

[0162] The comparison of the fuzzer FCIFuzz implemented based on this method with the control groups EPF and PAVFuzz in terms of the evaluation index of the number of triggered crashes is shown in Table 3. As can be seen from this table, on the target program Dnamasq, FCIFuzz triggered 20 more target program crashes on average than the fuzzer EPF, and 38.4 more fuzzer crashes than the fuzzer PAVFuzz. On the target program Live555, the fuzzer FCIFuzz implemented by this method triggered 35 more target program crashes on average than EPF, and 50.3 more fuzzer crashes on average than the fuzzer PAVFuzz. Finally, in terms of the total number of triggered target program crashes, the fuzzer FCIFuzz implemented by this method increased by 55 on average compared with the control group EPF, and 88.7 on average compared with the control group PAVFuzz. Therefore, the effectiveness of this method can be proved in terms of the number of triggered crashes index.

[0163] 4-2-3. Comparison of the number of generated seeds

[0164] Since there is no concept of seeds in the fuzzer PAVFuzz, the experimental results of the fuzzer FCIFuzz implemented based on this method and the control group EPF in terms of the number of generated seeds are shown in Table 4. As can be seen from this table, on the target program Dnsmasq, FCIFuzz generated 25.4 more seeds on average than the control group EPF, with a promotion rate of 16.57%. On the target program Live555, FCIFuzz generated 40.7 more seeds on average than the control group EPF, with a promotion rate of 11.49%. Finally, in terms of the total number of generated seeds, the fuzzer FCIFuzz implemented based on this method increased by 13.02% on average compared with the control group EPF. Therefore, the effectiveness of this method can be proved in terms of the number of generated seeds index.

[0165] Table 4 Results of the number of generated seeds

[0166]

[0167] 4-2-4. Experimental summary

[0168] Through experiments, overall, the fuzzer FCIFuzz implemented based on this method has certain improvements compared to the control group's fuzzers EPF and PAVFuzz, which are based on grammar generation and coverage information guidance, in terms of three evaluation metrics: the number of branch covers, the number of triggered crashes, and the number of generated seeds. Specifically, in terms of the total number of branch covers of the target program, FCIFuzz has an average of 236.9 more branch covers than EPF, representing a 7.66% increase; and an average of 326.9 more branch covers than PAVFuzz, representing a 10.88% increase. In terms of the total number of crashes of the target program triggered, FCIFuzz triggers an average of 55 more target program crashes than EPF; and an average of 88.7 more target program crashes than PAVFuzz. In terms of the total number of generated seeds, the fuzzer FCIFuzz implemented based on this method has an average increase of 13.02% compared to the control group EPF. Therefore, the effectiveness of the fuzzer FCIFuzz implemented based on this method can be proven through these three different experimental metrics.

Claims

1. A protocol vulnerability mining and testing method guided by fine-grained coverage information, characterized in that Including: A preprocessing stage, a test preparation stage, and a fuzz testing stage; Among them, the fuzz testing stage includes seed selection, test case generation, and test case evaluation; The test preparation stage includes the following sub-steps: 2-1. Build an execution engine; In order to record the coverage of each branch of the protocol entity program during the fuzz testing process, a global shared memory shareMem with a size of XM bytes is allocated when building the execution engine; In order to separately record the branch coverage of the current test case, a temporary shared memory temShareMem with the same size of XM bytes is allocated; During the fuzz testing process, if the current test case can be used as a new seed, the information stored in the temporary shared memory can be used as the branch coverage information of the seed for the protocol entity program; 2-2. Create the global data structures required during the fuzz testing process; In order to be able to count the number of times each branch is covered, a global integer array branCovArr is created. Each array element counts the number of times a branch is covered. To correspond to the global shared memory, the length of the integer array is taken as XM. Each element is an integer, which is 4 bytes, and the maximum value is 2 32 -1, which is sufficient to count the number of times the corresponding branch is covered during the fuzz testing process; the memory size occupied by the global integer array branCovArr is 4 * XM bytes; 2-3. System warm-up, initialize the data model seed library.

2. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 1, characterized in that Before the test preparation stage, there is also a preprocessing stage, including the following sub-steps: 1-1. Instrumentation and compilation of the protocol entity program source code; During the fuzz testing process, in order to obtain the branch coverage of the protocol entity program to be tested, the protocol entity program source code needs to be instrumented and compiled using the gcc compilation tool provided by aflfast to generate a binary executable program; 1-2. Define the data model set; By analyzing the request data packets corresponding to the protocol to be tested and combining with the protocol specification to be tested, use the data model definition function provided by the BooFuzz fuzz testing framework to define the data model set model of the protocol to be tested set ={model1,...,model i ,...,model n}, i = 1,..., n, where n is the total number of data models of the protocol to be tested, and the defined data model is used as the protocol specification template for generating corresponding test cases; 1-3. Define the session model; Combined with the requirements of the session message sequence for the protocol specification to be tested, use the session model definition function provided by the BooFuzz fuzz testing framework to connect the data model set model defined in steps 1-2 set in the data models into a session model sessionModels = {sessionSeq1,…,sessionSeq j ,…,sessionSeq m}, j = 1,..., m, where sessionSeq j represents a session sequence in the session model sessionModels, which consists of several data models with a sequential relationship, and m is the number of session sequences in the session model.

3. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 1, characterized in that The specific implementation of step 2-3, system warm-up, initialize the data model seed library, is as follows: 2-3-1. Traverse the session models sessionModels defined in steps 1-3, and sequentially traverse and extract the session sequences sessionSeq j , and proceed to step 2-3-2; if the traversal of the session models has ended, the system warm-up work has been completed; 2-3-2. According to the sequential relationship of the data model in the session sequence sessionSeq obtained in step 2-3-1 j Traverse the session sequence sessionSeq j The current data model model obtained i , go to step 2-3-3; if the traversal work for the session sequence sessionSeq j has been completed, then go to step 2-3-1; 2-3-3. According to the data model traversed in step 2-3-2 i Generate unmutated test cases and use them as data models i seedBank i The initial seed in i 0 , stands for seed bank seedBank i The 0th seed in the i 0 Input into the execution engine and create the branch coverage information array seedBranCovArr corresponding to the seed i 0 , traverse the temporary shared memory temShareMem, if the corresponding position of the temporary shared memory is not zero, that is, the branch represented by the corresponding position subscript is covered by the seed, add the branch represented by the subscript to the branch coverage information array seedBranCovArr of the seed i 0 And update the global integer array branCovArr to count the number of times each branch is covered; finally, the seed seed i 0 And its corresponding branch coverage information array seedBranCovArr i 0 Add to data model model i Corresponding seed bank seedBank i At this time, seedBank i Only {[seed i 0 ,seedBranCovArr i 0 ]}, and clear the temporary shared memory temShareMem; go to step 2-3-2 to initialize the seed library of the next data model.

4. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 1 or 3, characterized in that The fuzz testing stage includes the following sub-steps: 3-1. Data model selection; 3-2. Seed selection; 3-3. Test case generation; 3-4. Test case evaluation.

5. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 4, characterized in that The specific implementation of step 3-1 is as follows: 3-1-1. Traverse the session models sessionModels defined in steps 1-3, and sequentially traverse and extract the session sequences sessionSeq j , and go to step 3-1-2; if the traversal of the session model has ended, then end the fuzz testing work; 3-1-2. According to the sequential relationship of the data model in the session sequence sessionSeq obtained in step 3-1-1 j traverse the session sequence sessionSeq j The obtained current data model model i is used as the data model in the current fuzz testing phase and proceed to step 3-2; if the traversal of the session sequence sessionSeq j is completed, proceed to step 3-1-1.

6. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 5, characterized in that The specific implementation of step 3-2 is as follows: 3-2-1. Traverse the seed bank seedBank i , calculate the seed weight; If there is no seed in the seed bank i it means that the data model i has undergone sufficient fuzz testing at this stage, and proceed to step 3-1-2 to select a new data model as the data model to be used at this stage; If the seed bank seedBank i has only one seed, there is no need to sort the seeds, and directly go to step 3-2-3; otherwise, traverse the data model model selected in step 3-1-2 i corresponding seed bank seedBank i ={[seed i 0 , seedBranCovArr i 0 ,…,[seed i k , seedBranCovArr i k ,…,[seed i t , seedBranCovArr i t}, k = 1,..., t, where t is the number of seeds in the seed bank seedBank i Select a seed seed i k from it in turn, and calculate the weight of the seed seed i k according to the branch coverage information array seedBranCovArr i k of the seed seed i k . The branch coverage information array seedBranCovArr i k records which branches of the protocol entity program are covered by the seed seed i k ; The formula for calculating the weight of the seed is as follows: Among them, weight i k represents the weight of the seed i k seedBranCovArr i k is the branch coverage information array of the seed i k which records which branches are covered by the seed. The length of this array len(seedBranCovArr i k ) is the number of branches covered by the seed. Let r range from 0 to length(seedBranCovArr i k ) - 1 to traverse the array seedBranCovArr i k . Among them, branCovArr[seedBranCovArr i k [r]] is the number of times the branch represented by seedBranCovArr i k [r] is covered during the fuzz testing process. Let the number of times be w, then 1 / w 2 will rapidly approach 0 as w increases, that is, the rarer the branch covered by the seed, the greater the contribution of this branch to the weight i k value of the seed; After calculating the weight i k weight of the seed i k , store its weight information into the weight array seedWeight i corresponding to the seed bank seedBank i = [..., weight i k ,...]; Traverse the seed bank seedBank in the same way i Deposit the weight of each seed calculated during the process into the seed bank seedBank i The corresponding weight array seedWeight i , after the traversal of the seed bank is completed, obtain the seed weight array seedWeight i = [weight i 0 ,…,weight i k ,…,weight i t , k = 1,..., t, where t is the number of seeds in the seed bank seedBank i In, turn to step 3-2-2; 3-2-2. Sorting the seeds in the seed bank i in the seed bank; Using the seedWeight obtained in Step 3-2-1 i = [weight i 0 , …, weight i k , …, weight i t , sort the corresponding seeds in the seed bank seedBank i = {[seed i 0 , seedBranCovArr i 0 , …, [seed i k , seedBranCovArr i k , …, [seed i t , seedBranCovArr i t} in descending order of weight; 3-2-3. Select the seed; Select the first seed from the sorted seed bank i as the benchmark for generating test cases in the current stage, that is, the seed with the largest weight; assume the selected seed with the largest weight is seed i k ; and initialize the energy value of this seed, seedEnergy i k = SC, and the threshold of this seed is threshold i k , and use this seed, seed i k to execute the test case generation process.

7. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 6, wherein The specific implementation of step 3-3 is as follows: 3-3-1. Obtain the current seed i k The corresponding data model model i All variable fields in the seed i k The values in are fields = {field1,…,field p ,…,field q}, p = 1,..., q, where q is the number of variable fields in the data model model i Set a marker bit index for the currently mutated field position b And initialize the marker bit index b = 0; 3-3-2. Traverse all the values of the variable fields in the seed obtained in step 3-3-1, namely fields, and assume that the currently traversed field value is field p ; If, during the traversal of fields, the position of the last field in the seed has been reached, then reset the marker index b to 0 and start traversing again, so that the next round of mutation starts from the first unmutated field; 3-3-3. If the field value traversed in step 3-3-2 is field p It's a seed i k The value has not been mutated, and the field is in the mark position index b The following fields are then mutated using the EPF mutation engine. p Mutate to get the current field field p The mutation result of mutated_field p , and mark the index b Set to p; otherwise go to step 3-3-2 to continue traversing fields; 3-3-4. Use the current field p mutated_field obtained in step 3-3-3 p to replace the seed i k corresponding field value, and obtain the newly generated test case testcase; 3-3-5. According to the current session sequence sessionSeq j obtain the data model model i all the pre-data models, then generate the unmutated pre-message sequence preMessSeq according to the pre-data models, form the message sequence messSeq by combining the test case testcase generated in step 3-3-4 with the pre-message sequence preMessSeq, and inject the message sequence messSeq into the execution engine; go to step 3-4 for test case evaluation.

8. The protocol vulnerability mining and testing method guided by fine-grained coverage information according to claim 7, wherein The specific implementation of step 3-3 is as follows: 3-4-1. Monitor the status of the execution engine If the message sequence messSeq containing the new test case causes the execution engine to crash, a crash report is generated, and the test case testcase that triggers the crash is saved for post-fuzz testing analysis of the crash; 3-4-2. Evaluate whether a new branch is triggered Traverse the global shared memory shareMem to determine whether the currently executed test case testcase has triggered a new branch; If there are new non-zero bytes in the global shared memory shareMem, it means that a new branch is triggered; If a new branch is triggered, the current test case is used as a new seed, i x and an array of branch coverage information corresponding to this seed, seedBranCovArr, is created. i x Traverse the temporary shared memory temShareMem. If the value at the corresponding subscript position y in the temporary shared memory is not zero, that is, the branch represented by the subscript at the corresponding position is covered, add the branch represented by this subscript to the array of branch coverage information for this seed, seedBranCovArr. i x During the process of traversing the temporary shared memory temShareMem, update the global integer array branCovArr; if the value at the corresponding subscript position y in the temporary shared memory is not zero, increment the data value stored in branCovArr[y] by one to count and update the number of times each branch is covered; finally, add the seed i x and its corresponding array of branch coverage information, seedBranCovArr, i x to the current data model model i corresponding to the seed bank seedBank i and clear the temporary shared memory temShareMem. 3-4-3. Update and check the seed energy value Adjust the current stage template seed seed selected in step 3-2-3 according to the evaluation result of step 3-4-2 i k of the energy value seedEnergy i k ; if the evaluation result is to trigger a new branch, set the energy value seedEnergy i k to the initial value SC; Otherwise, seed i k Energy value of seedEnergy i k Set to seedEnergy i k ×ɑ, where ɑ is the attenuation factor, which is less than 1 and greater than 0. The smaller the value, the greater the attenuation of the seed energy value; After the update, if the energy value seedEnergy of the current template seed i k is less than i k <threshold i k , it means that the performance of the seed i k is not good. Delete the seed and the branch coverage information array of the seed from the seed bank i , and go to step 3-2 to select a new seed from the seed bank i as the current template seed; If the seed energy value seedEnergy i k >= threshold i k , then go to step 3-3-2 and continue to generate new test cases using this seed.

Citation Information

Patent Citations

  • Software hybrid fuzz testing method and equipment based on fine-grained information synchronization

    CN114036040A

  • Protocol software testing method and device

    CN115687158A