A vulnerability cloning detection system and method based on a binary tuple
Through a binary-based vulnerability cloning detection system, the problems of information redundancy and high detection false alarm rate in the prior art are solved, and more accurate and efficient vulnerability cloning detection are achieved.
Patent Information
- Application Number
- CN202210907348.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The existing vulnerability cloning detection methods based on code similarity have problems such as redundancy in information and high detection false positive rates, and it is difficult to accurately determine whether the detection result is a complete vulnerability.
Using a binary-based vulnerability cloning detection system, by generating a vulnerability feature library and performing abstract representation of the code attribute graph, vulnerability cloning detection is converted into a binary matching problem, simplifying the calculation process and improving detection efficiency.
It reduces information redundancy, reduces detection false alarm rate, and can accurately determine whether the detection result is a complete vulnerability, improving the accuracy and efficiency of vulnerability cloning detection.
Smart Images

Figure CN115270136B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a vulnerability cloning detection system and method based on a binary tuple, belonging to the technical field of vulnerability detection. Background Art
[0002] Vulnerability cloning detection of open-source software has always been one of the hot research issues in the field of software analysis. With the continuous advancement of Internet technology, network information security has become an issue that people pay more and more attention to. However, at the same time, the number of vulnerabilities exposed in the Internet also shows an increasing trend year by year. In addition, in the computer field, many enterprises have formed unique business models using open-source software. According to statistics, 99% of enterprises use open-source software in their IT systems. Open source has become the general trend of computer development today. While open-source software is becoming increasingly popular and growing, the number of vulnerabilities related to open-source software is also growing. This is mainly because when cloning and reusing open-source code, those vulnerable codes may be brought into the system under development. This leads to the hidden and widespread propagation of a vulnerability. From the perspective of attackers, open-source software security vulnerabilities are easily obtained from the Internet, which allows attackers to design targeted attacks based on the details of the official patches. It can be seen that understanding the vulnerability cloning of open-source software as much as possible can not only improve the security of the software itself, but also control the propagation of vulnerabilities during the reuse of open-source software code. Therefore, developing a vulnerability cloning detection system to quickly and accurately detect known vulnerabilities existing in the system is an effective measure to cope with software security risks and reduce various losses caused by vulnerabilities.
[0003] Currently, according to the analysis methods of vulnerabilities, software vulnerability detection can be divided into pattern matching-based methods and code similarity matching-based methods. Open-source software has natural convenience in obtaining source code, so this system mainly focuses on vulnerability detection of source code.
[0004] In early research, some algorithms sought to find unique patterns of specific types of vulnerabilities and find the parts in the code under test that match this pattern, that is, it was determined that there was a vulnerability. However, for this detection method, first of all, the generation of vulnerability patterns mainly relies on the experience of security experts, so it is highly subjective. For each type of jitter, different detection patterns need to be designed, and the design is complex. The imperfection of the rules will also lead to inaccurate detection. Another detection algorithm is based on the idea that code similar to a vulnerability is very likely to contain a vulnerability, and a method is designed to compare the similarity of codes.
[0005] Regarding the problem of vulnerability detection for open-source software, there are some code similarity-based methods to handle this problem. ReDeBug can quickly discover unpatched vulnerable code in an operating system-scale codebase. It uses a feature hashing method to encode n tokens in a bit vector, enabling ReDeBug to perform similarity detection in a very efficient manner. SecureSync uses an extended abstract syntax tree (xAST) to represent vulnerable code segments for reproducing vulnerability detection of copied source code. CBCD uses subgraph isomorphism matching to determine whether the PDG of defective code is a subgraph of the PDG of a software system and provides 4 optimization methods for PDG queries. Vuddy takes functions as the basic detection granularity and uniformly replaces the type names, variable names, and function names in vulnerable functions to ensure the detection of vulnerable code clones. Song et al. adopted the method of program slicing to extract statement blocks related to vulnerabilities for matching.
[0006] In existing source code vulnerability clone detection based on code similarity, since the abstract syntax tree, data flow graph, control flow graph, and program dependence graph are not comprehensive enough in abstracting code, the found vulnerability features are not complete. Therefore, the detection methods designed according to the above several schemes have low detection performance for vulnerabilities. The code property graph combines the above several abstract structures, so it can obtain enough code structure features. However, due to too much abstract information obtained, there are many redundancies, so the detection speed and accuracy are not high enough.
[0007] Another issue is that existing detection methods report on a function-by-function basis. However, a single vulnerability often involves multiple functions. Just because only some of the vulnerable functions are matched does not necessarily mean it is a vulnerability.
[0008] In summary, among current vulnerability clone detection methods, text-based methods only require lexical analysis of code, and the code information contained in the text is too little. Although the abstract syntax tree can obtain more code structure information than text representation, the structure of the abstract syntax tree is complex, and the cost of vulnerability detection for large-scale software is very high. The code structure information contained in the code property graph is more complete than that of the abstract syntax tree, but the calculation of graph similarity is also complex. The present invention first uses a binary tuple to represent the information in the code property graph, transforming the subgraph matching problem into a binary tuple matching problem, greatly simplifying the calculation process, improving the detection efficiency, and enabling it to be applied to larger-scale software.
[0009] In addition, the current vulnerability detection reports the detection results on a function-by-function basis. However, the cause of a vulnerability is complex, and fixing a vulnerability often involves multiple functions. When a CVE involves multiple functions, only some of the vulnerable functions included in the CVE may be matched in the project under test. Although some vulnerable code is included, the vulnerability of these functions has disappeared, that is, there is no vulnerability. Then the detection result of this function reported is meaningless. Therefore, we need to check and screen the detection results one by one to determine whether the reported results must contain a vulnerability. Therefore, for the case where a vulnerability involves multiple functions, in the detection process of this method, a vulnerability analysis and integration process oriented to CVE is added.
[0010] [1] JANG J, AGRAWAL A, BRUMLEY D. ReDeBug: finding un-patched code clones in entire OS distributions[C] / / 2012 IEEE Symposium on Security and Privacy (S&P). 2012: 48-62.
[0011] [2] LI Zan, BIAN Pan, SHI Wenchang, et al. Approach of leveraging patches to discover unknown vulnerabilities[J]. Journal of Software, 2018, 29(5): 1199-1212. LI Z, BIAN P, SHI W C, et al. Approach of leveraging patches to discover unknown vulnerabilities[J]. Journal of Software, 2018, 29(5): 1199-1212.
[0012] [3] Kim S, Woo S, Lee H, et al. VUDDY: A Scalable Approach for Vulnerable Code Clone Discovery[C] / / 2017 IEEE Symposium on Security and Privacy (SP). IEEE, 2017.
[0013] [4]Song X, Yu A, Yu H, et al. Program Slice Based Vulnerable Code Clone Detection[C] / / 2020 IEEE 19th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom). IEEE, 2020. Summary of the Invention
[0014] The technical problem to be solved by the present invention: Overcoming the deficiencies of the prior art, providing a vulnerability clone detection system and method based on a binary tuple, reducing information redundancy and at the same time reducing the false positive rate of detection; at the same time, further accurately judging whether the detection result is a complete vulnerability.
[0015] The technical solution of the present invention: A vulnerability clone detection system based on a binary tuple, including: a generation module for a vulnerability feature library, a vulnerability clone detection module, and a result filtering and generation module;
[0016] The generation module for the vulnerability feature library first generates code property graphs according to the vulnerable function and the patched function respectively, and then obtains the statements related to the vulnerability according to the code property graphs. After these statements are standardized and abstractly represented, vulnerability feature information is obtained, and finally a vulnerability feature library is formed;
[0017] In the vulnerability clone detection stage, the project to be tested is abstracted. First, the functions in the project to be tested are obtained, and then the functions are processed to generate code property graphs. The code property graphs are further abstractly represented. Finally, they are matched with the vulnerability feature library to obtain the detected vulnerabilities;
[0018] The result filtering and generation module uses a filter to further judge the detected vulnerabilities in units of CVE. If the result of the detected vulnerability contains all the vulnerable functions involved in a CVE-ID, it is determined to be a vulnerability the same as the CVE. If it contains some vulnerable functions, it is an uncertain result.
[0019] The specific implementation of the generation module for the vulnerability feature library is as follows:
[0020] (1) First, obtain the vulnerability function and the patching function corresponding to the vulnerability given by CVE (Common Vulnerabilities & Exposures); on the NVD website, for each CVE, its relevant resource URL will be published. Obtain the CVE dataset in json format on the NVD website, and then process the dataset. Since this invention is only used to detect C / C++ code, filter out the CVE information related to C / C++. Among the filtered CVEs, there are CVE numbers, patch URLs, and description information. According to the patch URL, obtain the corresponding patch and the vulnerability function. Modify the vulnerability function based on the patch and the vulnerability function, and finally obtain the patching function. After obtaining the vulnerability function and the patching function, mark the two functions respectively. After marking the two functions, process to obtain the code property graphs of the two functions;
[0021] (2) Process the obtained vulnerability function BadFunc and patching function GoodFunc to obtain the code property graphs of the two functions. Specifically: use joern to scan the functions, and after processing, the vulnerability function and the patching function both obtain two files, namely the node file and the edge file. The node file includes node keywords, node information, and node types. The edge file includes the edge types between two nodes, including eight types such as FLOW_TO, USE. According to these two files, abstract the two functions into two sets represented by binary tuples respectively. The definition of the binary tuple is as follows: [Code1_out, Code2_in], where the two statements in the binary tuple cannot be changed because the control flow or data flow relationship is implied therein;
[0022] (3) Obtain all relevant binary tuple representations of the statements with the marks '+' and '-' as the core in the vulnerability function BadFunc and the patching function GoodFunc. This process is called'slicing'. Integrate all the obtained binary tuples. The set of binary tuple representations obtained from the vulnerability function is called the vulnerability function slice set, and the set of binary tuple representations obtained from the patching function is called the patching function slice set. For these two sets, classify them again. The classification rules are as follows:
[0023] C c = Vul s ∩Pat s
[0024] B c = C c ∩Vul s
[0025] G c = C c ∩Pat s
[0026] Among them, C_c refers to the tuples that appear in both the set of slices of the vulnerable function and the set of slices of the patched function, B_c refers to the tuples that only appear in the set of slices of the vulnerable function, and G_c refers to the tuples that only appear in the set of slices of the patched function;
[0027] (4) Standardize the tuples in each set, represent the tuples as a 32-bit hash value. For each vulnerable function, the vulnerability features consist of the following five parts:
[0028] {CVE_ID#Funcname, Funchash, C_c_hash, B_c_hash, G_c_hash})
[0029] Among them, CVE_ID#funcname refers to the function name of the vulnerability, Funchash refers to the hash value of the function body after hashing the vulnerable function, C_c_hash refers to the hash value of the tuple after hashing C_c, B_c_hash refers to the hash value of the tuple after hashing B_c, and G_c_hash refers to the hash value of the tuple after hashing G_c;
[0030] (5) Finally, for the various features extracted for each vulnerable function, store them in the JSON data format. Organize the vulnerability features in units of functions. One function corresponds to one record, and store the finally obtained vulnerability feature information to form a vulnerability feature library.
[0031] The specific implementation of the vulnerability cloning detection stage module is as follows:
[0032] (1) First, process the project under test to obtain the functions under test in the project under test. The specific process of extracting functions includes: parsing and extracting the file name, function name, variable list in the function, parameter name list, data type list including user-defined variable types, function call list, and function body; save the above-parsed function content as a file in units of functions, and the file name naming format is as follows: file path#~file name$~function name$function range in the file;
[0033] (2) After extracting and saving the function, abstract the function to generate a code property graph. Then compress the information of the code property graph. The code property graph is a directed graph. The nodes in the graph contain code statements, and the edges contain the control and dependency relationships between the codes. According to the order of the nodes in the code property graph, two code statements connected by the same directed edge are used as elements of a binary tuple. Abstract the binary tuple, and then standardize the binary tuple. Use the hash algorithm - FNV-1a to hash the binary tuple to generate a binary tuple hash, which is convenient for subsequent detection operations;
[0034] (3) The final detection process is to compare and match the hash generated by abstracting the function to be tested with the hashes in the vulnerability feature library. The specific implementation process is as follows: First, compare the hash of the function to be tested with the C_c_hash in the vulnerability feature library. By comparing, determine whether the function to be tested is relevant to a specific vulnerability. If it is relevant to a specific vulnerability, then it is necessary to further determine whether the function to be tested is a vulnerable function; If the hash of the function to be tested is more in line with the set of binary tuples B_c_hash of the unique features of the vulnerability, and the degree of compliance with the set of binary tuples G_c_hash of the features of the patched function is not high, then determine that the function is a vulnerable function. Finally, the function to be tested will be marked as a vulnerable function, and the tested project contains this vulnerability. The function to be tested will be put into the vulnerability list.
[0035] The specific implementation in the result filtering and generating module is as follows:
[0036] (1) Re-organize all the vulnerable functions collected in units of CVE to generate the JSON file cve_write.json. The data organization form in the file is as follows:
[0037] {CVE_ID:[CVE_ID#funcname1, CVE_ID#funcname2,...]};
[0038] CVE_ID refers to the vulnerability number officially released by CVE. CVE_ID#funcname1 refers to the function name involved in this CVE. The reorganization of the content in the vulnerability feature library is for the convenience of subsequent result filtering;
[0039] (2)Finally, compare the detection results obtained in the vulnerability cloning detection stage with the above-mentioned files. Treat the file content as a set and calculate the inclusion relationship between sets. If all functions under the CVE_ID are included in the detection results, it is determined that there must be a vulnerability given by this CVE_ID in the project under test. Report this vulnerability as a confirmed vulnerability in the project under test to the security personnel, and the security personnel can directly repair it according to the patch given by the official vulnerability website. Otherwise, it cannot be directly determined that there is such a vulnerability in the project under test, and these detected functions will be reported as suspected vulnerabilities.
[0040] A method for detecting vulnerability cloning based on a binary tuple according to the present invention comprises the following steps:
[0041] (1) First is the generation part of the vulnerability feature library. Generate code property graphs for the vulnerability function and the patched function respectively to obtain the vulnerability function code property graph and the patched function code property graph. With the vulnerability marking statements given by the official vulnerability website as the center, according to the statement relationships in the property graphs, obtain the subgraphs of the related statements of the vulnerability statement code in the two property graphs respectively. The specific implementation is as follows: First, according to the released patch, find the statements directly related to the vulnerability marked in the patch, and then, with these statements as the center points, find the subgraphs generated by the statements related to these statements in the vulnerability function code property graph. The main purpose of this step is to reduce the interfering statements in the code, that is, the statements irrelevant to the vulnerability. Then, abstract the subgraphs of these two vulnerability function code property graphs into binary tuples. The content of the binary tuple is two statements with an order relationship in the directed graph. This step further compresses the subgraphs to reduce redundant information, so that the complexity during matching is much smaller than graph matching. Calculate the hash values of these binary tuples using the hash function, and divide these binary tuple hashes into three sets, namely, the statement slice binary tuple hash set related to the deleted statements, the statement slice binary tuple hash set related to the added statements, and the statement slice binary tuple hash set related to both the deleted and added statements. The main purpose of this step is to facilitate the screening of functions with vulnerable code, then the screening of functions that may be vulnerabilities, and then the confirmation of whether it is a vulnerable function during the detection process;
[0042] (2) Then in the detection process, the functions in the project under test are first extracted. Then, for the function under test, a code attribute graph is first generated. Then, all the statement node pairs in the code attribute graph are abstracted into tuples. Then, the tuples are hashed to obtain the corresponding hash values. For all the tuple hashes generated by the function under test, the statement slice tuple hash set related to both the deletion and addition statements is first matched. If the set threshold 1 is reached, the statement tuple hash set related to the deletion statement in the vulnerability feature library is then matched. If the set threshold 2 is reached, the statement slice tuple hash set related to the addition statement is finally matched. If the threshold 3 is met, the detection process is terminated and it is determined that the function under test is a vulnerable function. If all three matching conditions cannot be met, the function under test is not a vulnerable function. Finally, a function under test that matches the vulnerability feature in the vulnerability library is obtained. These functions are the vulnerabilities detected in the project under test. The values of the three thresholds need to be determined through system experiments. The threshold determination rules are as follows: First, for threshold 1, in the experiment, we must ensure that more than 99% of the vulnerability-related functions in the experimental data set can be screened out as candidate functions. For threshold 2, in the experiment, we must ensure that 99% of the suspected vulnerability functions are screened out from the candidate functions screened by threshold 1. For threshold 3, in the experiment, it was found that as long as it is <1-threshold 2, it can have a good detection distinction between the vulnerability function and the patched function.
[0043] (3) Analyze and filter the detection results. Since a vulnerability often involves multiple files or functions, after detecting the project in units of functions and obtaining the vulnerable functions, it is necessary to re-analyze and filter from the overall perspective of the vulnerability. The direct result report in units of functions can only determine that there are clones of the vulnerable function code in the tested project, but whether the cloned vulnerable code can cause a vulnerability needs further manual confirmation. Design a result analysis filter. This filter uses CVE_ID as a unit and saves multiple function names involved in a CVE in the form of key-value pairs for result analysis. Finally, all analysis results are organized in the form of CVE. If the detection result contains all functions involved in CVE_ID, it is directly determined to be a vulnerability with the same CVE_ID, and the report will give a list of confirmed results. If it contains some, it is an uncertain result, and the report will give a list of suspected results. In this way, the confirmed results directly determined as vulnerabilities can directly refer to the patch link provided by the CVE official website for vulnerability repair. Only those suspected vulnerability results need to be analyzed to confirm whether they need to be repaired.
[0044] The advantages of the present invention compared with the prior art are:
[0045] (1) The present invention proposes a new vulnerability feature representation method based on binary tuples, generating code property graphs for the vulnerable function and the patched function respectively. Considering these two generated abstract graphs, first, according to the official patch, find the statements marked in the patch that are directly related to the vulnerability, and then, taking these statements as the center points, find the statement fragments in the property graph that are related to these statements. The main purpose of this step is to reduce the interfering statements in the code, that is, the statements unrelated to the vulnerability. Next, abstract the subgraphs of these two code property graphs into binary tuples. The content in the binary tuple is mainly two statements with an order relationship in the directed graph. This step mainly further compresses the subgraph to reduce redundant information. In this way, the complexity during matching is much less than graph matching. Subsequently, divide these binary tuples into three sets, namely the set of statements related to deleted statements, the set of statements related to added statements, and the set of statements related to both deleted and added statements. The main purpose of this step is to first screen out the functions with vulnerable code during the detection process, then screen out the functions that may be vulnerable, and finally confirm whether it is a vulnerable function. Finally, comprehensively analyze the detection results and design a filter. Taking CVE as the unit, if all vulnerable functions are included in the results, then it is determined to be a vulnerability identical to a specific CVE. If only some are included, it is an uncertain result, which can further compress the obtained abstract graph information and reduce the complexity of matching.
[0046] (2) The present invention proposes a result filtering method for CVE, which can further analyze and summarize the detection results and reduce the manual workload. Brief Description of the Drawings
[0047] Figure 1 is the overall block diagram of the system of the present invention;
[0048] Figure 2 is the implementation flowchart of the vulnerability clone detection module of the present invention;
[0049] Figure 3 is the processing process of the function to be tested. Detailed Embodiment
[0050] The present invention will be described in detail below with reference to the drawings and embodiments.
[0051] As Figure 1As shown in the figure, the system of the present invention includes three major modules: a vulnerability feature library generation module, a vulnerability clone detection module, and a result filtering and generation module. The vulnerability feature extraction module obtains statements related to vulnerabilities from the vulnerable functions and patched functions. After standardization and abstract representation, these statements become the features of specific vulnerabilities; in the vulnerability clone detection stage, the project to be tested is abstracted and matched with the vulnerability feature library. The result analysis report refers to further analyzing these detected vulnerabilities in units of CVE.
[0052] The implementation steps of the system's vulnerability feature library generation module, vulnerability clone detection module, and result filtering and generation module will be introduced in detail below.
[0053] 1. Vulnerability feature library generation module
[0054] First, obtain the vulnerable functions and patched functions corresponding to the CVE. On the NVD website, for each CVE, its relevant resource website will be published. First, obtain the CVE dataset in json format on the NVD website, and then process the dataset to filter out the CVE information related to C / C++. The filtered CVE contains the CVE number, patch website, description information, etc. According to the patch website, obtain the corresponding patch and vulnerable function. The patched function cannot be obtained directly. According to the patch and the vulnerable function, modify the vulnerable function to finally obtain the patched function. After obtaining the vulnerable function and the patched function, it is necessary to re-mark the statements marked with "-" in the patch to the vulnerable function and the statements marked with "+" to the patched function respectively.
[0055] After the two functions are marked, it is also necessary to process to obtain the code property graph of the function. Therefore, when marking, the method cannot be the same as that marked in the patched function. Therefore, "-" is marked as " / / -" at the end of the statement, and the same treatment is done for "+".
[0056] The second part is to process the obtained vulnerable function (BadFunc) and patched function (GoodFunc) to obtain their code property graphs. Use joern to scan the functions, and then each function will get two files after processing, a node file and an edge file. The node file includes node keywords, node information, node types, etc. The edge file includes the edge types between two nodes, including eight types such as FLOW_TO and USE. According to these two files, the function is abstracted into a binary tuple representation. The binary tuple is defined as follows: [Code1_out,Code2_in], where the two statements in the binary tuple cannot be changed. Because the control flow or data flow relationship is implied in it.
[0057] The third part is in BadFunc or GoodFunc. According to the obtained binary tuples, obtain the statements with " / / -" or " / / +" as the core. The slicing rules are as follows:
[0058] 1) First, for the marked statements, directly perform forward and backward slicing. Forward slicing means that when slicing, find the marked code and the code at the position of Code2 in the binary tuple, and put Code1 into the forward slice set. Backward slicing means that when slicing, find the code at the position of Code1 of the marked code, put Code2 into the backward slice set, and put this binary tuple into the vulnerability feature set.
[0059] 2) For the forward slice set, perform forward slicing and put the binary tuple into the vulnerability feature set.
[0060] 3) For the backward slice set, perform backward slicing and put the binary tuple into the vulnerability feature set.
[0061] Integrate all the obtained binary tuples, mainly the set of binary tuples generated by the " / / -" statements and their slices. It is called the vulnerability function slice set (Vulnerability slices, Vul_s), and the set of binary tuples generated by the " / / +" statements and their slices is called the patch function slice set (Patch slices, Pat_s).
[0062] For these two sets, classify them again, and the definitions are as follows:
[0063] C c = Vul s ∩ Pat s
[0064] B c = C c ∩ Vul s
[0065] G c = C c ∩ Pat s
[0066] Among them, C_c refers to the binary tuples that appear in both the vulnerability function slice set and the patch function slice set. B_c refers to the binary tuples that only appear in the vulnerability function slice set. G_c refers to the binary tuples that only appear in the patch function slice set.
[0067] The fourth part is to standardize the binary tuples in each set and represent them as a 32-bit hash value.
[0068] Finally, for each vulnerability function, the vulnerability features consist of the following four parts:
[0069] {CVE_ID#Funcname,Funchash,C_c_hash,B_c_hash,G_c_hash}
[0070] 2. Vulnerability cloning detection module, such as Figure 2 shown.
[0071] The entire detection module consists of two main parts, namely the source code preprocessing part for processing the project under test and the vulnerability detection part.
[0072] First is the processing of the project under test, mainly including two steps:
[0073] 1) Obtain functions by processing files. The files in the project under test need to be processed to obtain a data processing set with functions as units. The main process of processing is to traverse all files, find the files ending with ".C", ".Cpp", ".C*", and extract the functions in these files respectively.
[0074] 2) Obtain an abstract representation by processing functions. The above-extracted functions are analyzed using Joern to obtain their code property graphs. Then, based on the code property graphs, all statement pairs in the functions are obtained to form a set of function binary tuples. Abstract and standardize the binary tuples to obtain a code representation that can match the vulnerability feature library. The specific process of processing the function under test is as Figure 3 shown. First, use Joern to analyze the function to obtain two files, the Edges file and the Nodes file. Among them, the Edges file stores the edge attribute information in the code property graph, including nodes and the edges between nodes. The Nodes file stores the node attribute and content information, and gives the corresponding information according to the node number. The information in the Nodes file and the Edges file can be connected through the node number. After obtaining the code property graph file, the two statement nodes connected by the edges in the code property graph can be extracted to form a statement binary tuple (code1, code2). Standardize these binary tuple statement pairs, mainly by uniformly replacing variable names, parameter names, and type names with "VAR", "PARA", "TYPE". Finally, calculate the hash value for the standardized binary tuples, and each binary tuple gets a hash value.
[0075] Then is the vulnerability detection part in the detection module, and the main process is:
[0076] First, during the matching process of the function under test, it is necessary to match three sets respectively, and set the similarity thresholds for matching different sets. First, match the target code with the common context set. When a certain threshold is reached, confirm that the target code is a vulnerability-sensitive function. Match it with the vulnerability feature set. When a certain threshold is reached, it is found that there may be a vulnerability. Finally, match it with the patched feature set. If the threshold is not exceeded, confirm that the target code is a vulnerable function rather than a function after patching.
[0077] 3. Result Filtering and Generation Module
[0078] The reason for a vulnerability to form is complex, and the repair of a vulnerability often involves multiple functions. When a CVE involves multiple functions, it may be possible to match some of the vulnerable functions, but the vulnerability of these functions has disappeared. Therefore, for the case where a vulnerability involves multiple functions, during the detection process, a vulnerability analysis and integration process oriented to CVE is added. If a vulnerability involves multiple functions, if all the functions of a vulnerability are detected in the target project, then it can be determined that it is a vulnerability. If only some of the vulnerable functions are detected, it is impossible to determine whether it is a vulnerability, so a determination result of a suspected vulnerability is given.
[0079] Finally, two files are given. One file is the definite vulnerability detection result, and the other is the suspected vulnerability detection result. In the files, taking CVE as a unit, the function under test, the name of the vulnerable function, the vulnerable function, and the corresponding patch file are displayed.
[0080] Although the specific implementation methods of the present invention are described above, those skilled in the art should understand that these are only examples. Without departing from the principles and implementation of the present invention, various changes or modifications can be made to these implementation schemes. Therefore, the protection scope of the present invention is defined by the appended claims.
Claims
1. A vulnerability cloning detection method based on a binary tuple, characterized in that: (1) Generate a vulnerability feature library. Generate code property graphs for the vulnerability function and the patched function respectively to obtain the vulnerability function code property graph and the patched function code property graph. Taking the vulnerability marking statement given by the official vulnerability website as the center, according to the statement relationships in the property graph, obtain the subgraphs of the related statements of the vulnerability statement code in the two property graphs respectively, and then abstract the subgraphs of these two vulnerability function code property graphs into binary tuples. The specific implementation is as follows: First, according to the released patch, find the statements directly related to the vulnerability marked in the patch, and then take these statements as the center points to find the subgraphs generated by the statements related to these statements in the vulnerability function code property graph. Then abstract the subgraphs of these two vulnerability function code property graphs into binary tuples. The content in the binary tuple is two statements with an order relationship. Perform hash calculation on these binary tuples using a hash function to obtain the corresponding hash values. Divide these binary tuple hashes into three sets, namely the statement slice binary tuple hash set related to deletion statements, the statement slice binary tuple hash set related to addition statements, and the statement slice binary tuple hash set related to both deletion and addition statements; (2) During the detection process, first extract the functions to be tested in the project to be tested to generate a code property graph, then abstract all statement node pairs in the code property graph into binary tuples, and perform hash calculation on the binary tuples to obtain the corresponding hash values. For all binary tuple hashes generated by the functions to be tested, first match the statement slice binary tuple hash set related to both deletion and addition statements. If the set reaches the set threshold 1, then continue to match the statement slice binary tuple hash set related to deletion statements in the vulnerability feature library. If the set reaches the set threshold 2, then finally match the statement slice binary tuple hash set related to addition statements. If the threshold 3 is satisfied, the detection process ends, and it is determined that the function to be tested is a vulnerability function. If all three matching conditions are not satisfied, then the function to be tested is not a vulnerability function. Finally, obtain a function to be tested that matches the vulnerability features in the vulnerability library, and these functions to be tested are the detected vulnerabilities that appear in the project to be tested; (3) Analyze and filter the detection results. Since a vulnerability often involves multiple functions, after detecting the project by function unit and obtaining the vulnerability functions, analyze and screen from the overall vulnerability perspective again. The direct result report in function unit can only determine the cloning of the vulnerability function code in the project to be tested, but whether the cloned vulnerability code can cause a vulnerability needs further analysis. Design a result analysis filter. This filter takes the CVE_ID as the unit and saves the names of multiple functions involved in a CVE in the form of key-value pairs for result analysis. All analysis results are organized in the form of CVE; If all the functions involved in the CVE_ID are included in the detection result, it is directly determined that there is a vulnerability identical to the CVE_ID, and the report will give a list of definite results; if only some are included, it is an uncertain result, and the report will give a list of suspected results.
2. A vulnerability clone detection system based on a binary tuple, which is used to implement the method described in claim 1. Characterized in that: It includes: A generation module for a vulnerability feature library, a vulnerability clone detection module, and a result filtering and generation module; The generation module for the vulnerability feature library first generates code property graphs based on the vulnerable function and the patched function respectively, and then, according to the code property graphs, obtains the statements related to the vulnerability. After these statements are standardized and abstractly represented, vulnerability feature information is obtained, and finally a vulnerability feature library is formed. In the vulnerability clone detection stage, the project to be tested is abstracted. First, the functions in the project to be tested are obtained, then the functions are processed to generate code property graphs, the code property graphs are further abstractly represented, and finally they are matched with the vulnerability feature library to obtain the detected vulnerabilities. The result filtering and generation module uses a filter to further judge the detected vulnerabilities in units of CVE. If all the vulnerable functions involved in a vulnerability number CVE_ID are included in the result of the detected vulnerability, it is determined that there is a vulnerability identical to the CVE_ID in the project to be tested; if only some vulnerable functions are included, it is an uncertain result.
3. The vulnerability clone detection system based on a binary tuple according to claim 2. Characterized in that: The specific implementation of the generation module for the vulnerability feature library is as follows: (1) First, obtain the vulnerable function and the patched function corresponding to the vulnerability given by the Common Vulnerabilities and Exposures (CVE) of the general vulnerability disclosure; on the NVD website, for each CVE, its relevant resource website is published. Obtain the CVE data set in json format on the NVD website, and then process the data set to screen out the CVE information; the screened CVE information includes the CVE number, the patch website, and the description information; according to the patch website, obtain the corresponding patch and the vulnerable function; according to the patch and the vulnerable function, modify the vulnerable function, and finally obtain the patched function; after obtaining the vulnerable function and the patched function, mark the two functions respectively. After the two functions are marked, process to obtain the code property graphs of the two functions. (2)Process the obtained vulnerable function BadFunc and patched function GoodFunc to obtain the code property graphs of the two functions. Specifically: Use joern to scan the functions. After processing, both the vulnerable function and the patched function obtain two files, namely the node file and the edge file. The node file includes node keywords, node information, and node types. The edge file includes the edge types between two nodes, including eight types such as FLOW_TO and USE. According to these two files, abstract the two functions into two sets represented by binary tuples respectively. The definition of the binary tuple is as follows: [Code1_out,Code2_in]. The order of the two statements in the binary tuple cannot be changed because it implies its control flow or data flow relationship. (3)Obtain all relevant binary tuple representations of the statements with the markers '+' and '-' as the core in the vulnerable function BadFunc and the patched function GoodFunc. This process is called slicing. Integrate all the obtained binary tuples. The set of binary tuple representations obtained from the vulnerable function is called the vulnerable function slice set, and the set of binary tuple representations obtained from the patched function is called the patched function slice set. Classify these two sets again. The classification rules are as follows: Among them, C_c refers to the binary tuples that appear in both the vulnerable function slice set and the patched function slice set. B_c refers to the binary tuples that only appear in the vulnerable function slice set. G_c refers to the binary tuples that only appear in the patched function slice set. (4)Standardize the binary tuples in each set and represent the binary tuples as a 32-bit hash value. For each vulnerable function, the vulnerability features consist of the following five parts: {CVE_ID#Funcname, Funchash, C_c_hash, B_c_hash, G_c_hash} Among them, CVE_ID#funcname refers to the function name of the CVE vulnerability. Funchash refers to the hash value of the function body after hashing the vulnerable function. C_c_hash refers to the hash value of the binary tuple after hashing C_c. B_c_hash refers to the hash value of the binary tuple after hashing B_c. G_c_hash refers to the hash value of the binary tuple after hashing G_c. (5)Finally, store the various features extracted for each vulnerable function in JSON data format. Organize the vulnerability features in units of functions. One function corresponds to one record, and store the finally obtained vulnerability feature information to form a vulnerability feature library.
4. The binary tuple-based vulnerability cloning detection system according to claim 2, characterized in that: The specific implementation of the vulnerability cloning detection phase module is as follows: (1)Firstly, the item to be measured is processed to obtain the function to be measured in the item to be measured. The process of extracting the function to be measured includes: parsing and extracting the file name, function name, variable list in the function, parameter name list, data type list, function call list, and function body; saving the parsed and extracted function to be measured as a file in units of functions, and the file name naming format is as follows: file path#~file name$~function name$range of the function in the file; (2)After the function to be measured is extracted and saved, it is abstracted to generate a code property graph, and then the information of the code property graph is compressed. The code property graph is a directed graph, where the nodes in the graph contain code statements, and the edges contain the control and dependency relationships between the codes. According to the order of the nodes in the code property graph, two code statements connected by the same directed edge are used as elements of a binary tuple, the binary tuple is abstracted, and then the binary tuple is standardized. The hash algorithm is used to calculate the hash of the binary tuple to generate a binary tuple hash; (3)The final detection process is to compare and match the hash generated by abstracting the function to be measured with the hash in the vulnerability feature library. The specific process is: firstly, compare the hash of the function to be measured with the C_c_hash in the vulnerability feature library to determine whether the function to be measured is relevant to a certain vulnerability through the comparison. If it is relevant to a certain vulnerability, it is necessary to further determine whether the function to be measured is a vulnerable function; if the hash of the function to be measured is more consistent with the binary tuple set B_c_hash of the unique vulnerability features and less consistent with the binary tuple set G_c_hash of the patch function features, then it is determined that the function is a vulnerable function; finally, the function to be measured will be marked as a vulnerable function, the item to be measured contains this vulnerability, and the function to be measured will be put into the vulnerability list.
5. The binary tuple-based vulnerability cloning detection system according to claim 2, characterized in that: The specific implementation in the result filtering and generating module is as follows: (1)All the vulnerable functions collected are reorganized in units of CVE to generate a JSON file, and the organization form of the file content is as follows: {CVE_ID : [ CVE_ID#funcname1, CVE_ID#funcname2, ……]}; CVE_ID refers to the vulnerability number officially released by CVE, and CVE_ID#funcname1 refers to the function name involved in this CVE. The reorganization of the content in the vulnerability feature library is for the convenience of subsequent result filtering; (2) Compare the detection results obtained in the vulnerability cloning detection phase with the JSON file. Treat the content of the JSON file as a set and calculate the inclusion relationship between sets. If all functions under the CVE_ID are included in the detection results, it is determined that there is a vulnerability given by this CVE_ID in the project under test. Report this vulnerability as a vulnerability that is definitely present in the project under test to the security personnel, and the security personnel can directly repair it according to the patch given on the official vulnerability website. Otherwise, it cannot be directly determined that there is such a vulnerability in the project under test, and these detected functions will be reported as suspected vulnerabilities.
Citation Information
Patent Citations
Fine-grained source code vulnerability detection method based on graph neural network
CN111259394A
Graph-based source code vulnerability detection system
US20210279338A1