Test Case Generation Method for Vulnerability Root Cause Location
By inserting the basic block weights in the program and exploring the paths in stages, and combining the dynamic energy allocation algorithm to generate test cases, the path exploration generated by test cases in the existing technology is solved, and the accuracy and efficiency of vulnerability root cause positioning is improved.
Patent Information
- Application Number
- CN202211123212.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-09-15
AI Technical Summary
When generating test cases, the existing technology cannot effectively constrain path exploration, resulting in overfitting the statistical results of spectrum-based vulnerability positioning technology or being unable to analyze the correct results, and lack of high-quality test cases, which affects the accuracy of vulnerability root cause positioning.
By identifying the weight of the basic block based on the crash path, we explore the path in stages and use the dynamic energy allocation algorithm to generate test cases, improve the similarity and convergence of path exploration, and generate high-quality test cases.
It significantly improves the quality of test cases, reduces interference items that are not related to the source of the vulnerability, and improves the accuracy and efficiency of the source of the vulnerability.
Smart Images

Figure CN115422070B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a test case generation method for locating the root cause of vulnerabilities. Background Art
[0002] The improvement of software testing technology has made the discovery of program crashes a completely automated process. However, the rapid growth in the number of crashes and the increasing scale and complexity of software have made it a time-consuming, costly, and difficult task to manually locate and fix vulnerabilities. The long vulnerability exposure period has led to a sharp increase in the probability of vulnerabilities being exploited, resulting in high-risk consequences and even economic losses or loss of life. Research shows that software maintainers spend half or more of their time just locating the root cause of vulnerabilities. Therefore, the vulnerability repair work has a very high demand for automated vulnerability location technology, which can help software maintainers effectively locate the root cause of vulnerabilities with minimal manual intervention.
[0003] The spectrum-based fault localization technology (SFL) is considered to be one of the most prominent technologies in this area of research due to its efficiency and effectiveness. According to statistical analysis, among the papers published since 1977, the research on spectrum-based fault localization related to spectrum technology accounts for 42%. Its working principle is to execute the given test cases and assign a probability estimate or score to each element of the program (such as statements, blocks, functions) according to the results of executing the test cases, and evaluate the possibility of a vulnerability being triggered when this statement is executed. A high score indicates that executing this statement is sufficient and necessary to trigger the vulnerability, and it is most likely that patching here by software maintainers will eliminate the vulnerability.
[0004] The research focus in this area lies in improving the statement scoring algorithm or complex combinations of single predicates. However, the statistical analysis method alone is not sufficient to produce high-fidelity results because it is not known in which input to statistically calculate the probability scores. Almost all spectrum-based vulnerability location studies assume the existence of appropriate test inputs, but this is not the case in reality. High-quality test cases (a collection of test inputs) are an essential part of this technology. Experiments have shown that without well-designed test cases, it often leads to overfitting of statistical results or the inability to analyze the correct results due to the confounding nature of test inputs. This problem is very serious in large programs where more than 1000 program statements can be observed in a single execution. Summary of the Invention
[0005] In view of the fact that existing research focuses on improving the scoring algorithm of statements or complex combinations of single predicates, but the statistical analysis method alone is not sufficient to produce high-fidelity results. Because it is not known in which input to statistically calculate the probability score, and randomly generated test cases often lead to overfitting of statistical results, or it is impossible to analyze the correct results due to the complexity of test inputs. A test case generation method for locating the root cause of vulnerabilities is proposed. According to three steps: information instrumentation, seed scheduling, and phased path exploration, good path constraints are imposed on the generation of test cases, and the dynamic energy allocation algorithm is used to balance the relationship between the divergence and convergence of path exploration, so that the generated test cases neither contain confounding terms unrelated to the crash path nor overfit to the original crash path.
[0006] To achieve the above object, the present invention inserts corresponding code into different basic blocks of the original program through information instrumentation based on the crash path to identify the weights of the basic blocks. Then, according to the instrumentation information, corresponding scores are assigned to the initial seeds and divided into different categories. Finally, path exploration is carried out in stages according to the crash path, and test cases are generated in the final stage. The present invention is constrained by the crash path, improving the quality of the generated test cases, greatly reducing the possibility of the appearance of interference terms unrelated to locating the root cause of vulnerabilities, and using the energy scheduling algorithm to well balance the problem of divergence and convergence when exploring paths, making the generated test cases more conducive to accurately locating the root cause of vulnerabilities.
[0007] A test case generation method for locating the root cause of vulnerabilities, comprising:
[0008] Through information instrumentation based on the crash path, corresponding code is inserted into different basic blocks of the original binary program to identify the weights of the basic blocks; in order to maintain the similarity of the explored paths, we need to mark the paths associated with the crash execution in the program, and then diverge and explore the remaining similar paths according to this path; here, the selection of the instrumentation granularity is at the basic block level. When instrumenting, it is mainly judged whether there are functions and function sequences related to the original crash path in the basic block, and different weights are assigned to the basic blocks based on this.
[0009] According to the instrumentation information, corresponding scores are assigned to the initial seeds and divided into different categories, and seed scheduling is carried out based on dynamic energy allocation;
[0010] For seeds of different categories, path exploration is carried out in stages according to the crash path, and test cases are generated in the final stage.
[0011] Further, inserting corresponding code into different basic blocks of the original binary program to identify the weights of the basic blocks includes:
[0012] Determine whether there are functions and function sequences related to the original crash path in the basic block, and assign different weights to the basic blocks based on this.
[0013] Furthermore, the determining whether there are functions and function sequences related to the original crash path in the basic block and assigning different weights to the basic blocks based on the determination includes:
[0014] Determine whether there is a Call Funcname instruction in the current basic block, where Funcname is the function name in the stack traceback function list List at the time of the crash. Then assign a score Score to the basic block according to the number of corresponding instructions in the current basic block. If there is no Call Funcname instruction in the current basic block, it is assigned to the initial value.
[0015] Determine the position of Funcname in this basic block in List and record it as Order. If there are multiple Funcnames, first select the function sequence with the largest number of occurrences as the Order of the basic block. If the number of occurrences is the same, select the first Funcname that appears as Order. If there is no function in List in the basic block, all are assigned the initial value.
[0016] The obtained scores and orders corresponding to each basic block are injected into each basic block, and the scores corresponding to each basic block are used as the weights of each basic block.
[0017] Furthermore, assigning corresponding scores to the initial seeds according to the plugging information and dividing them into different categories, and performing seed scheduling based on dynamic energy allocation includes:
[0018] First, the states of the seeds are divided according to their characteristics, including: the initial run of the program after the stub is input, that is, the seed, and then the seed is given a corresponding score s_score according to the score of the basic blocks through which the seed passes, and the value of the Order label that appears most in the basic blocks through which the seed passes is assigned to the seed sequence label s_order, and the initial value is assigned if there is no Order label. The seeds with a comprehensive score exceeding the set threshold are defined as guided seeds (vectoring). In the path exploration stage, guided seeds are mainly responsible for making the detection process quickly approach the predetermined target, and the rest of the seeds are ordinary seeds;
[0019] Then, the simulated annealing algorithm model is used for dynamic energy allocation. Initially, it is guided by probability. As the number of seeds selected increases, it gradually transitions to being guided by seed scores, achieving the goal of balancing divergent exploration and convergent exploration.
[0020] Furthermore, for seeds of different categories, path exploration is performed in stages according to the crash path, and test cases are generated in the final stage, including:
[0021] Put the guiding seeds with the same s_order into the same seed queue Order_seed_queue, and enable the corresponding Order_seed_queue during the corresponding Order exploration stage. Seeds with a score lower than the set threshold are put into the Order1 seed sequence; each Order seed sequence represents an exploration stage.
[0022] When exploring the path, assign energy to the seeds according to dynamic energy distribution. First, enter the divergent exploration mode, and gradually turn into the convergent exploration mode as the seed energy changes. After the exploration of the current stage is completed, merge the current seed queue with the previous seed queue.
[0023] In the path exploration of the last stage, save the generated crash test cases and non-crash test cases to form the final test cases.
[0024] Further, for different types of seeds, conduct path exploration in stages according to the crash path. After generating test cases in the last stage, it also includes:
[0025] Use the generated test cases as input and execute them again in the original binary program, and trace the execution information. Finally, obtain the location of the vulnerability root cause based on the statistical analysis of the execution information.
[0026] Compared with the prior art, the beneficial effects of the present invention are:
[0027] The present invention first constrains the path exploration and guides the entire exploration process based on the crash path. Then, the seeds are divided into guiding seeds and ordinary seeds, and path exploration is carried out in stages according to different types of seeds, greatly improving the efficiency of approaching the target point. Then, by introducing a dynamic energy regulation algorithm to dynamically allocate the seed energy, the problem of divergence and convergence during path exploration is well balanced, greatly improving the quality of the generated test cases, so as to better improve the accuracy of vulnerability root cause location. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic flowchart of a test case generation method for vulnerability root cause location according to an embodiment of the present invention;
[0029] Figure 2 It is an example diagram of staged path search according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following further explains the present invention with reference to the drawings and specific embodiments:
[0031] As Figure 1 shown, a test case generation method for vulnerability root cause location (abbreviated as Dloc) includes:
[0032] First, Dloc instruments a given target binary program to obtain a binary file with basic block weight score information. The Dgenerate (test case generation) component divides the seeds (i.e., the inputs of the software) into categories through the given original inputs and the instrumented binary file to generate different seed queues, and then generates test cases through steps such as dynamic energy allocation and mutation. Dloc takes the generated test cases as inputs and executes them again in the original binary program, and tracks the execution-time information. Finally, the location of the vulnerability source is obtained based on the statistical analysis of the execution-time information.
[0033] The specific introduction is as follows:
[0034] 1 Information instrumentation based on crash path guidance
[0035] To maintain the similarity of the exploration paths, we need to mark the paths associated with the crash execution in the original program, and then diverge and explore the remaining similar paths based on this path. Here, the selection of our instrumentation granularity is the basic block level. Since a basic block has only one exit and all instructions within the block will be executed, it is reasonable to select the basic block as the smallest granularity.
[0036] Algorithm 1 Information instrumentation algorithm based on sensitive functions
[0037]
[0038] Algorithm 1 describes our instrumentation process. The present invention introduces two new global variables, Score and Order. After the instrumented program executes the input, Dgenerate can read these values as the basis for subsequent seed type judgment and mutation strategy selection. Specifically, during the instrumentation process, we insert corresponding code for different basic blocks to identify the weights of the basic blocks, and like the coverage information of the traditional fuzz testing tool AFL, it is globally shared in the way of shared memory.
[0039] List refers to the stack backtrace function list at the time of crash, which records the function names and function call sequences called when the program crashes. The IfHasFunc function in line 8 of the algorithm is used to determine whether there is an instruction like CallFuncname in the current basic block, where Funcname is the function name in the List list. Then, the basic block is assigned a score (Score) according to the number of corresponding instructions in the current basic block, and if not, it is uniformly assigned an initial value.
[0040] The CheckOrder function on line 9 of the algorithm is used to determine the position of Funcname in the List sequence in this basic block and record it as Order. If there are multiple Funcnames, first select the function sequence with the most occurrences as the basic block Order. If the number of occurrences is the same, select the first occurring Funcname sequence as Order. If there is no function in the List list in the basic block, it is uniformly assigned an initial value. We need to execute a similar execution path to the crash path, and the same function call sequence reflects the similarity of the execution trajectory.
[0041] 2 Seed Scheduling Strategy Based on Dynamic Energy Allocation
[0042] The seed energy determines the number of times a seed is selected during the path exploration process. To better generate test cases, we propose a seed scheduling strategy based on dynamic energy allocation.
[0043] First, we divide the seeds into different states according to their characteristics. Run the instrumented program with an initial input once, and then give each seed a corresponding score s_score according to the score of the basic blocks passed by the seed. And assign the value that appears most frequently for the Order label in the basic blocks passed by the seed to the s_order of the seed sequence label. If there is none, assign the initial value.
[0044] Algorithm 2: Dynamic Energy Allocation Algorithm
[0045]
[0046] The variable H determines the switching of the exploration mode at each stage, and its initial value is 0.5. T is the temperature in the simulated annealing algorithm, and x.times is the number of times the seed is selected. As the number of selections increases, the value of T will gradually decrease. The energy K of the initial seed ori Complies with the energy assignment rule of the original AFL, that is, it is allocated according to the coverage information. At this time, it belongs to the divergent exploration mode. As the number of times the seed is selected increases, H gets closer and closer to its score s_score, and the seed energy K also changes accordingly, switching from the coverage-guided scheduling strategy to the seed score-guided scheduling strategy. At this time, it belongs to the convergent exploration mode, so as to achieve the switching of the exploration mode.
[0047] 3 Phased Path Exploration Strategy Based on Seed Categories
[0048] In 2, we assigned the seed a seed score s_score and a seed sequence label s_order, and based on this, proposed a phased path exploration strategy based on seed categories. First, the seeds whose comprehensive scores exceed the set threshold are defined as vectoring seeds. In the path exploration stage, vectoring seeds are mainly responsible for making the exploration process quickly approach the predetermined target. In order to prevent the basic structure from being destroyed by the mutation strategy, vectoring seeds have a mutation strategy different from that of ordinary seeds. AFL predefines eleven mutation operations, and its mutation scheduling strategy is divided into three different stages, namely the deterministic stage, the destruction stage (Havoc stage), and the splicing stage (Splicing stage). Because vectoring seeds mainly play the role of directional path exploration, the mutation operators in the Havoc and Splicing stages are likely to destroy their original structure and thus affect the effect of their directional guidance. In order to ensure that their original structure is not excessively destroyed, we only retain the mutation operators in the Deterministic stage for this type of seeds. Then put the guiding seeds with the same s_order into the same seed queue Order_seed_queue, and enable the corresponding Order_seed_queue seed pair in the corresponding Order exploration phase. Seeds with scores lower than the set threshold will be placed in the Order1 sequence.
[0049] Algorithm 3: Phased path exploration algorithm
[0050]
[0051] Each Order seed sequence represents an exploration phase, such as Figure 2 As shown, taking Order2 as an example, at this stage we only focus on exploring the program path between the function HintFile and the function gf_hinter_track_process. First, enable the guided seed in the Order2 sequence and merge it with the previous seed queue. When exploring the path, enter the divergent exploration mode first. The goal at this time is to explore as many program paths as possible. Seeds that cover more basic blocks are given more energy. As the number of times the seed is selected increases, the T value will gradually decrease, and the corresponding seed energy K will become closer and closer to the original s_score score. At this time, enter the convergent exploration mode, the purpose is to make the divergent path closer to the target function. After the current stage of exploration is completed, the Order2 seed sequence will be merged into the Order3 seed sequence. In the last stage of path exploration, the generated crash test cases and non-crash test cases will be saved to form the final test cases.
[0052] In summary, the present invention first constrains the path exploration and guides the entire exploration process based on the crash path. Then, the seeds are divided into guiding seeds and ordinary seeds, and the path exploration is carried out in stages according to different types of seeds, greatly improving the efficiency of approaching the target point. Then, by introducing a dynamic energy regulation algorithm to dynamically allocate the energy of the seeds, the problem of divergence and convergence during path exploration is well balanced, greatly improving the quality of the generated test cases, so as to better improve the accuracy of locating the root cause of vulnerabilities.
[0053] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A test case generation method for vulnerability root cause location, characterized in that, Including: By performing information instrumentation based on the crash path, inserting corresponding code for different basic blocks in the original binary program to identify the weights of the basic blocks; The step of, by performing information instrumentation based on the crash path, inserting corresponding code for different basic blocks in the original binary program to identify the weights of the basic blocks includes: judging whether there are functions and function sequences related to the original crash path in the basic block, and assigning different weights to the basic blocks based on this; the step of judging whether there are functions and function sequences related to the original crash path in the basic block and assigning different weights to the basic blocks based on this includes: determining whether there is a Call Funcname instruction in the current basic block, where Funcname is the function name in the function list List of the stack backtrace at the time of the crash, and then assigning a score Score to the basic block according to the number of corresponding instructions in the current basic block, and assigning a unified initial value if not; determining the position of Funcname in this basic block in List and recording it as Order, if there are multiple Funcnames, first selecting the function sequence with the most occurrences as the Order of the basic block, if the number of occurrences is the same, selecting the first-occurring Funcname as the Order, and if there is no function in List in the basic block, assigning a unified initial value; injecting the obtained Score and Order corresponding to each basic block into each basic block, and using the Score corresponding to each basic block as the weight of each basic block; Assigning corresponding scores to the initial seeds according to the instrumentation information and classifying them into different categories, and performing seed scheduling based on dynamic energy allocation; For seeds of different categories, performing path exploration in stages according to the crash path, and generating test cases in the final stage; the step of, for seeds of different categories, performing path exploration in stages according to the crash path and generating test cases in the final stage includes: putting the leading seeds with the same seed sequence label s_order into the same seed queue Order_seed_queue, enabling the corresponding Order_seed_queue in the corresponding Order exploration stage, and putting the seeds with scores lower than the set threshold into the Order1 seed sequence; each Order seed sequence represents an exploration stage; when exploring the path, assigning energy to the seeds according to dynamic energy allocation, first entering the divergent exploration mode, and gradually turning into the convergent exploration mode as the energy of the seeds changes. After the exploration of the current stage is completed, merging the current seed queue with the previous seed queue; saving the generated crash test cases and non-crash test cases in the path exploration of the last stage to form the final test cases.
2. The test case generation method for vulnerability root cause location according to claim 1, characterized in that The step of assigning corresponding scores to the initial seeds according to the instrumentation information and classifying them into different categories, and performing seed scheduling based on dynamic energy allocation includes: First, the states of the seeds are classified according to their characteristics, including: running the input, i.e., the seeds, through the instrumented program once initially, then assigning a corresponding score s_score to the seeds according to the scores of the basic blocks passed by the seeds, and assigning the value that appears most frequently for the Order label in the basic blocks passed by the seeds to the seed sequence label s_order. If there is none, an initial value is assigned. Seeds with a comprehensive score exceeding the set threshold are defined as guiding seeds, and the remaining seeds are ordinary seeds; Next, the simulated annealing algorithm model is used for dynamic energy allocation. Initially, it is guided by probability, and gradually transitions to being guided by the seed scores as the number of times the seeds are selected increases, achieving the purpose of balancing divergent exploration and convergent exploration.
3. The test case generation method for vulnerability root cause localization according to claim 1, characterized in that For different types of seeds, path exploration is carried out in stages according to the crash path. After generating test cases in the final stage, it also includes: Regarding the generated test cases as input and re-executing them in the original binary program, tracking the information during execution, and finally obtaining the location of the vulnerability root cause based on the statistical analysis of the information during execution.
Citation Information
Patent Citations
Mode-based dynamic vulnerability discovery integrated system and mode-based dynamic vulnerability discovery integrated method
CN104598383A
Python software fuzzy test method based on dynamic type perception
CN110399300A