Code vulnerability detection method and vulnerability detection system based on graph simplification
Through a graph-based simplification method, program dependency graphs are simplified, weighted images are generated, and vulnerability detection is used to solve the problem of low detection efficiency and accuracy in the prior art, and more efficient and accurate vulnerability detection is achieved.
Patent Information
- Application Number
- CN202510010515.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing code vulnerability detection methods based on graph neural networks have low detection efficiency and accuracy, making it difficult to effectively extract vulnerability features, and are affected by redundant information and complex graph structures.
A graph simplification-based method is adopted to generate weighted images through preprocessing, graph simplification, vector embedding and central analysis, and a binary classification task is used to complete vulnerability detection.
It significantly improves the efficiency and accuracy of vulnerability detection, reduces the impact of redundant information, improves the accurate learning of vulnerability patterns, and reduces detection time.
Smart Images

Figure CN119939600A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of code vulnerability detection, and in particular to a code vulnerability detection method and a vulnerability detection system based on graph simplification. Background Art
[0002] In the context of open source software, the dependencies between programs / software are getting stronger and stronger. Complex software is often composed of multiple open source software, which are combined and interdependent. Together with the maintainers and developers who contribute to each open source software, they constitute the open source software supply chain. The development method using open source components will inevitably spread high-risk vulnerabilities to other components or projects through the dependencies between components, leading to further expansion of the danger. In recent years, network security incidents have occurred frequently, such as network attacks and user information leaks, so the security of open source software plays a vital role in the security of the entire software supply chain.
[0003] Vulnerability detection is an important issue in software supply chain security, but the diversity and complexity of the syntax and semantics of real-world vulnerability codes make it difficult to effectively extract vulnerability features when analyzing large programs. In addition, real-world vulnerability codes contain a large amount of redundant information that is irrelevant to the vulnerability, which will further aggravate the above problems. Traditional detection methods are limited by predefined rules, while existing vulnerability detection models based on graph neural networks usually retain all code information to extract vulnerability features. Since the proportion of vulnerability code in actual functions is very small, the rich information in the composite graph that is irrelevant to any vulnerability pattern hinders the accurate learning of vulnerability patterns. And due to the large number of nodes in the graph, vulnerability detection based on graph neural network models (GNN) is usually very time-consuming, resulting in low efficiency in detecting vulnerabilities in existing technologies. Summary of the invention
[0004] Based on this, it is necessary to provide a code vulnerability detection method and vulnerability detection system based on graph simplification to address the problems of low detection efficiency and accuracy of existing vulnerability methods based on graph neural networks.
[0005] In a first aspect, the present invention provides a code vulnerability detection method based on graph simplification, which comprises the following steps:
[0006] S1, preprocessing the original software code code0 to remove noise, and obtaining the preprocessed software code code1; wherein code1 includes K1 lines of code;
[0007] Perform image generation processing on code1 to obtain the program dependency graph PDG;
[0008] Among them, PDG includes K1 original nodes and K2 original node edges;
[0009] The k1th original node represents the k1th line of code in code1; k1∈[1, K1];
[0010] K2 original node edges represent the dependency relationship between K1 original nodes;
[0011] S2, perform graph simplification on PDG to obtain a program simplified graph CSG;
[0012] CSG includes N1 simplified nodes and N2 simplified node edges; N1<K1, N2<K2;
[0013] The graph simplification processing method includes:
[0014] S21, merging adjacent original nodes in the PDG according to the node type, the merged simplified nodes inherit the original node edges of the original nodes before the merger, and form simplified node edges;
[0015] S22, remove the noise nodes and free nodes in the PDG, and delete the original node edges connected to the noise nodes and free nodes;
[0016] S3, embed the N1 simplified nodes in the CSG into vectors to obtain N1 simplified node vectors, and update the CSG into a code vector graph;
[0017] S4, performing centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image;
[0018] S5 adjusts the weight image to the input specification requirements of the CNN model and performs a binary classification task to complete the vulnerability detection of the software code.
[0019] As a preferred example, in S21, the method for merging adjacent original nodes according to node type includes the following steps:
[0020] Determine whether the adjacent original nodes meet the merging rules;
[0021] If the adjacent original nodes meet the merging rules, the adjacent original nodes are merged;
[0022] The merging rules include:
[0023] If the parent node type is an expression statement and the child node type is also an expression statement, they will be merged;
[0024] If the parent node type is an identifier declaration statement and the child node type is an identifier, they are merged;
[0025] If the parent node type is a conditional statement, all child nodes of any type will be merged;
[0026] If the parent node type is a for loop statement and the child node type is a parameter list, they will be merged;
[0027] If the parent node type is a function call statement and the child node type is an identifier or parameter list, they are merged;
[0028] The method of merging adjacent original nodes includes retaining the original node of the parent node type, deleting the original node of the child node type, and retaining the original node edge corresponding to the original node of the child node type;
[0029] Among the adjacent original nodes, the parent node is the original node located in front, and the child node is the original node located in the back.
[0030] As a preferred example, in S22, the method for determining the noise node includes the following steps:
[0031] S221, determining the attributes of the predefined noise node;
[0032] S222, comparing the attributes of the K1 original nodes and the predefined noise nodes in the PDG one by one;
[0033] If the attribute of an original node is consistent with the attribute of a predefined noise node, the original node is determined to be a noise node.
[0034] As a preferred example, the method for determining the predefined noise node includes the following steps:
[0035] S2211, preprocessing the sample software code to remove noise, and performing image generation processing on the preprocessed sample software code to obtain a program dependency graph PDG';
[0036] Among them, the sample software code is software code that is known to have vulnerabilities;
[0037] S2212, counting the node types in PDG';
[0038] S2213, conduct an ablation experiment on each node type of PDG', and regard the node types that have a negative impact on the detection results as suspicious node types;
[0039] S2214, performing an ablation experiment on each node in the suspicious node type, and determining the nodes that have a negative impact on the detection result as noise nodes;
[0040] The properties of the noise node are recorded and saved as a predefined noise node.
[0041] As a preferred example, in S22, the method for determining a free node includes the following steps:
[0042] S2215, calculating the degrees of K1 original nodes in the PDG;
[0043] S2216, the original node with a degree of 0 is regarded as a free node.
[0044] As a preferred example, in S4, the method for generating a weight image comprises the following steps:
[0045] S41, calculate the destructibility metrics x1~x1 of N1 simplified nodes in the code vector graph N1 , eigenvector centrality EC1~EC N1 and closeness centrality CC1~CC N1 ;
[0046] S42, based on the destructibility metrics x1~x1 of N1 simplified nodes N1 The destructibility measurement channel is calculated;
[0047] According to the eigenvector centrality of N1 simplified nodes EC1~EC N1 Calculate the characteristic vector channel;
[0048] According to the closeness centrality of N1 simplified nodes CC1~CC N1 The approach channel is calculated;
[0049] S43, generating a weight image according to the destructibility metric channel, the feature vector channel, and the proximity channel;
[0050] The weight image is used to characterize the importance of the N1 row simplification nodes.
[0051] As a preferred embodiment, the method for calculating the destructibility measurement channel includes the following steps:
[0052] S411, calculate the destructibility measure x of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is:
[0053]
[0054] Where n reachable Indicates the number of simplified node edges between the n1th simplified node and its adjacent simplified nodes, n total Represents the number of all simplified nodes in the code vector graph;
[0055] S412, multiplying the n1th simplified node vector by the destructibility measure x of the n1th simplified node n1 Get channel value one;
[0056] S413, traverse N1 simplified nodes to obtain N1 channel values 1;
[0057] Arrange the N1 channel values in the order of code lines to obtain the destructibility measurement channel.
[0058] As a preferred embodiment, the method for calculating the feature vector channel comprises the following steps:
[0059] S421, calculate the eigenvector centrality EC of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is:
[0060]
[0061] Where α represents the proportional constant, and β represents the eigenvector centrality EC of the n1th simplified node. n1 The initial value of; A represents the adjacency matrix; A n1j1 represents the edge weight or connection relationship between the n1th simplified node and the j1th simplified node; x j1 Represents the eigenvector centrality EC of the j1th simplified node j1 , the j1th simplified node and the n1th simplified node are adjacent nodes, j1∈[1,N1];
[0062] S422, multiply the vector of the n1th simplified node by the eigenvector centrality EC of the n1th simplified node n1 Get channel value 2;
[0063] S423, traverse N1 simplified nodes to obtain N1 channel values 2;
[0064] Arrange the N1 channel values in the order of code lines to obtain the feature vector channel;
[0065] The method for calculating the approach channel includes the following steps:
[0066] S431, calculate the closeness centrality CC of the n1th simplified node n1 , and its calculation formula is:
[0067]
[0068] Where, d n1j2 represents the shortest distance between the n1th simplified node and the j2th node; j2∈[1, N1]; N represents the number of all simplified nodes in CSG;
[0069] S432, multiply the vector of the n1th simplified node by the closeness centrality CC of the n1th simplified node n1 Get channel value three;
[0070] S433, traverse N1 simplified nodes to obtain N1 channel values three;
[0071] Arrange the N1 channel values in the order of code lines to obtain the approximate channel.
[0072] As a preferred example, in S5, the method for adjusting the weight image to the CNN model input specification requirements includes the following steps:
[0073] If the size of the weight image is less than the threshold, zero vectors are filled at the end of the destructibility metric channel, feature vector channel, and proximity channel;
[0074] If the size of the weight image is larger than the threshold, the destructibility metric channel, the feature vector channel, and the end data of the near channel are deleted.
[0075] In a second aspect, the present invention further proposes a code vulnerability detection system based on graph simplification, which uses the code vulnerability detection method based on graph simplification in the first aspect. The code vulnerability detection system based on graph simplification includes: a preprocessing module, a graph simplification module, a node embedding module, an image generation module and a classification module.
[0076] The preprocessing module is used to preprocess the original software code code0 to remove noise, and obtain the preprocessed software code code1; it is also used to perform image generation processing on the preprocessed software code to obtain the program dependency graph PDG.
[0077] The graph simplification module is used to simplify the PDG to obtain the program simplified graph CSG.
[0078] The node embedding module is used to perform vector embedding on the N1 simplified nodes in the CSG, obtain N1 simplified node vectors, and update the CSG into a code vector graph.
[0079] The image generation module is used to perform centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image.
[0080] The classification module is used to adjust the weight image to the input specification requirements of the CNN model and perform binary classification tasks.
[0081] The beneficial effects of the present invention are:
[0082] 1. The present invention designs a new code vulnerability detection method based on graph simplification, which can represent vulnerable code as a graph while eliminating redundant information irrelevant to the vulnerability as much as possible, and then performs centrality analysis on the graph after removing redundant information to obtain a weight image with the importance of all code lines, and captures relevant features in the weight image through the CNN model to complete the vulnerability detection task. According to experimental tests, the vulnerability detection efficiency and accuracy of the code detected by the present invention are significantly higher than some existing code vulnerability detection methods, providing a more effective and feasible solution strategy for vulnerability detection in large-scale projects.
[0083] 2. The present invention simplifies the graph by reducing the size of the program dependency graph to reduce the distance between nodes and deleting free nodes and noise nodes. The simplified graph can reduce other noise interference that is not related to the vulnerability while retaining the structural features in the code, thereby improving the accuracy of vulnerability identification and reducing vulnerability detection time. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0085] Figure 1 A step diagram of a code vulnerability detection method based on graph simplification in an embodiment;
[0086] Figure 2 Flow chart of the code vulnerability detection method based on graph simplification in the embodiment;
[0087] Figure 3 A data graph comparing the detection efficiency of the code vulnerability detection method based on graph simplification provided by the present invention and the existing vulnerability detection methods. DETAILED DESCRIPTION
[0088] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0089] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0090] See also Figure 1 and Figure 2 This embodiment provides a code vulnerability detection method based on graph simplification, which includes the following steps:
[0091] S1. Preprocess the original software code code0 to remove noise. The preprocessed software code code1 includes K1 lines of code. Perform image generation processing on code1 to obtain a program dependency graph PDG. PDG includes K1 original nodes and K2 original node edges. The k1th original node represents the k1th line of code in code1; k1∈[1, K1]. The original node edges represent the dependency relationship between the original nodes.
[0092] In this step, preprocessing the software code code0 can reduce the noise information therein. Specifically, the first step is to remove the comments and blank lines in the software code code0, because these contents are usually irrelevant to the actual semantics of the program. The second step is to map the user-defined variables to unified symbolic names (for example, map all variables to "VAR1", "VAR2", etc.). The last step is to map the user-defined function names to standardized symbolic names, such as mapping functions to "FUN1", "FUN2", etc. Through such a normalization process, it is ensured that when extracting the program dependency graph PDG of the software code code1, the focus is on its structure and logical function, rather than specific naming or format details. Among them, the open source tool Joern is used to generate the program dependency graph PDG for the preprocessed software code code1, which is a graphical representation of the software code. The program dependency graph contains information about data flow and control in the code, and each node in the graph represents one of the lines of code data in the software code code1.
[0093] S2. Simplify the PDG to obtain a program simplified graph CSG. The CSG includes N1 simplified nodes and N2 simplified node edges; N1<K1, N2<K2.
[0094] It is well known that vulnerability detection methods based on graph neural network models (GNN) in the prior art have limitations in handling long-distance connections between nodes that are not directly adjacent to each other. Although stacking multiple layers of GNNs can be used to learn the global information of the graph (i.e., the long-term dependencies between nodes), the over-smoothing problem caused by deep GNNs will lead to similar embeddings of nodes with different labels, thereby reducing the overall performance. Existing vulnerability detection methods still have difficulties in capturing the global information of source code, which hinders their performance in detecting software vulnerabilities, especially for complex graphs.
[0095] In this step, the graph simplification process is achieved through two aspects, including:
[0096] S21. Merge adjacent original nodes in the PDG according to the node type. The merged simplified nodes inherit the original node edges of the original nodes before the merge and form simplified node edges, that is, reduce the distance between nodes by reducing the size of the program dependency graph PDG.
[0097] S22, remove the noise nodes and free nodes in the PDG, and delete the original node edges connected to the noise nodes and free nodes.
[0098] The program dependency graph PDG in the present invention aims to compress repeated information into the graph through type-based graph simplification processing, which can reduce the size of the graph and the distance between nodes.
[0099] Specifically, in S21, the method for reducing the size of the program dependency graph includes merging adjacent original nodes according to node types. Specifically, the method for merging adjacent original nodes according to node types includes the following steps:
[0100] Determine whether the adjacent original nodes meet the merging rules;
[0101] If the adjacent original nodes meet the merging rules, the adjacent original nodes are merged.
[0102] In this embodiment, five types of graph simplification merging rules are determined, as shown in Table 1.
[0103] Table 1: Merge rules in type-based graph simplification
[0104] rule Parent node type Child node type 1 Expression Statements Expression Statements 2 Identifier declaration statement Identifier 3 Conditional Statements Any type after the parent node 4 for loop statement Parameter List 5 Function call statement Identifier, parameter list
[0105] Among them, in a pair of adjacent original nodes, the parent node is the original node located in front of the original node edge, and the child node is the original node located behind the original node edge. The front and back are determined according to the control dependency direction of the original node edge. For a pair of adjacent original nodes that match a merge rule, the merge content includes retaining the original node of the parent node type, deleting the original node of the child node type, and retaining the original node edge of the original node of the child node type. The original node of the child node type will be deleted because its information is a refinement of its parent node and can also be reflected in the nodes behind it. For example, according to the merge rule in Table 1, the information in the child node "(malloc,malloc(50*sizeof(char)))" is replaced by the parent node "( <operator>.assignment,VAR1=(char*)malloc(50*sizeof(char)))", the child node is merged with the parent node and the child node is deleted. Since the parent node and the child node are "merged two by two", but are still on one path (edge). If the noise node is removed alone, the path (edge) between the A node in front of the noise node and the B node behind the noise node may be directly disconnected and disconnected. In this way, if there are other connected paths (edges) between the A and B nodes, the distance between the A and B nodes will be longer. Therefore, merging adjacent nodes according to node type can reduce the distance between nodes from the side.
[0106] In S22, some original nodes in the program dependency graph PDG obtained in S1 are irrelevant to understanding and learning the vulnerability patterns in the code, and are defined as "noise nodes". For example, the original node ("method", FUN1) only defines the name of a function, but it is connected to many other original nodes, which may cause interference when extracting feature information from other original nodes. To this end, it is necessary to identify and exclude the interference of these noise nodes as much as possible. Among them, the method for determining whether an original node in the program dependency graph PDG is a noise node includes the following steps:
[0107] S221, determining the attributes of the predefined noise node;
[0108] S222, compare the attributes of the K1 original nodes and the predefined noise nodes in the program dependency graph PDG one by one:
[0109] If the attribute of the original node in the program dependency graph PDG is consistent with the attribute of the predefined noise node, then the original node in the program dependency graph PDG is determined to be a noise node.
[0110] Among them, the predefined noise node as a judgment benchmark plays a key role. For this reason, it is important to determine the predefined noise node. The method for determining the predefined noise node includes the following steps:
[0111] S2211, preprocessing the sample software code to remove noise, and performing image generation processing on the preprocessed sample software code to obtain a program dependency graph PDG'. The sample software code is a software code known to have vulnerabilities.
[0112] S2212. Count the node types in PDG'.
[0113] In order to identify the types of noise nodes in the program dependency graph PDG', we first counted the number of all node types in the program dependency graph PDG', and classified each node in the graph according to its attributes or functions. The main node types can be divided into identifiers, arithmetic and logical operations, declaration statements, function calls, literals, unknowns, etc.
[0114] S2213. Perform an ablation experiment on each node type of PDG' and regard the node types that have a negative impact on the detection results as suspicious node types.
[0115] In this step, ablation experiments are performed on the node types that are counted. By comparing the detection results of retaining and removing these node types, the node types that have a negative impact on the detection results are judged as suspicious node types. This step preliminarily determines the possible noise node types.
[0116] S2214: Perform an ablation experiment on each node in the suspicious node type, and determine the nodes that have a negative impact on the detection result as noise nodes. Record the attributes of the noise nodes and save them as predefined noise nodes.
[0117] This step conducts a separate removal experiment for each node under the suspicious node type. Through the ablation experiment, the specific impact of removing each node on the overall detection performance is detected, and the nodes that have a negative impact on the detection results are judged as noise nodes. The properties of these noise nodes are recorded and saved as predefined noise nodes, and the predefined noise nodes are divided into five types:
[0118] "METHOD": declare the function name
[0119] "METHOD_RETURN": the return value of the main function
[0120] "RETURN": return value
[0121] "LITERAL": literal
[0122] "UNKNOWN": unknown type
[0123] The existence of these predefined noise nodes in the program dependency graph PDG' interferes with the detection results, and removing them can improve the accuracy of detection.
[0124] In the whole process of determining the predefined noise nodes, the Sent2vec model is selected for node vector embedding, and the TextCNN model is uniformly used to extract features in the detection result reading stage. The comparison of the impact of different types of nodes on the results is shown in Table 2.
[0125] Table 2: Comparison of different types of nodes
[0126]
[0127]
[0128] The superscript a in Table 2 indicates that no noise nodes are removed; the superscript b indicates that all noise nodes are removed. It can be seen from Table 2 that: removing non-noise nodes FIELD_IDENTIFIER or <op>.indirectFieldAccess (FIELD_IDENTIFIER is a field identifier; indirectFieldAccess is an indirect field access) has a significant impact on the detection results, compared to PDG without removing noise nodes a For the detection accuracy (Acc), it results in a decrease of about 3%. On the other hand, removing the predefined noise nodes METHOD, METHOD_RETURN, and RETURN will have a positive impact on the accuracy, thereby improving the detection accuracy. These predefined noise nodes are connected to many other nodes and may cause interference when other nodes extract feature information. For several other types of predefined noise nodes, removing them individually has little effect on the overall detection accuracy. However, when all types of predefined noise nodes are removed, the detection results are optimal. This shows that comprehensive consideration and removal of all predefined noise nodes can optimize the performance of the detection model and improve the final detection accuracy.
[0129] In addition, the free nodes in the program dependency graph PDG are judged by the size of the degree centrality, and the judgment method includes the following steps:
[0130] S215. Calculate the degrees of K1 original nodes in the program dependency graph PDG.
[0131] S216. Nodes with a degree of 0 represent nodes without edges, which are generally called isolated nodes. This type of node has little impact on the entire software code, but will increase the difficulty of vulnerability detection. Therefore, the degrees of K1 original nodes in the program dependency graph PDG are traversed, and the degrees of original nodes with a degree of 0 are regarded as free nodes and removed.
[0132] CFG (control flow graph), PDG (program dependency graph) and CPG (code property graph) are used as the benchmarks for program representation on two datasets, QEMU and FFmpeg. CSG in the present invention can achieve better results. The above four program representations are compared in experiments, and Sent2vec and CodeBERT are used as node embedding methods. In the model readout stage, TextCNN is uniformly used to extract features. The results are shown in Table 3.
[0133] Table 3: Vulnerability detection results using different program representations and node embedding methods
[0134]
[0135]
[0136] As can be seen from Table 3, the detection effect of using PDG as code representation is slightly better than CFG. Although CFG can describe control logic such as branches and loops in the program structure, it does not consider the correlation between data, which makes it limited in detecting vulnerabilities including data flow and variable operations. PDG's more comprehensive information makes it possible to more accurately capture possible defects in the code. In contrast, CPG provides more information about code properties, covering aspects such as grammatical structure, data flow, and control flow. However, the inclusion of a lot of irrelevant information may interfere with the model's learning of vulnerability patterns, and this effect is exacerbated when CodeBERT is used as a semantic extractor. The CSG proposed in the present invention combines graph-based representation and graph simplification to delete some nodes that are not related to the vulnerability. At the same time, it can be seen from Table 3 that CSG did achieve better results in the experiment.
[0137] S3. Perform vector embedding on the N1 simplified nodes in the CSG to obtain N1 simplified node vectors, and update the CSG into a code vector graph.
[0138] In this step, the pre-trained CodeBERT model is selected for vector embedding. The CodeBERT model consists of a bidirectional transformer, which can capture long-range code sequence dependencies. In addition, it can effectively maintain the relationship between contexts, collect potential fragile code patterns, and minimize information loss. The CodeBERT model used at the same time also adopts a multi-attention head structure, which enables the model to focus on multiple key points of a code sequence. At this point, CSG is transformed into a new form, the so-called code vector graph. Each node in the code vector graph is embedded with a vector.
[0139] S4. Perform centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image.
[0140] The purpose of this step is to effectively convert the code vector graph into a weighted image that takes into account the contribution of different code lines to the program semantics. To this end, the code vector graph is considered as a social network, and social network centrality analysis is applied to obtain the importance of all code lines. This embodiment uses three types of centrality to calculate the importance of N1 simplified nodes, including the destructibility measure x n1 , Eigenvector centrality EC n1 and closeness centrality CC n1 . It uses the destructibility measure x1~x N1 Calculate the destructibility measurement channel; use the eigenvector centrality EC1~EC N1 Calculate the eigenvector channel; use the proximity centrality CC1~CC N1 Calculate the approach channel. The calculation methods of the three channels are as follows:
[0141] 1. The method for calculating the destructibility metric channel includes the following steps:
[0142] S411. Calculate the destructibility measure x of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is:
[0143]
[0144] Where n reachable Indicates the number of simplified node edges between the n1th simplified node and its subsequent adjacent simplified nodes, ignoring the simplified node edges between the n1th simplified node and its preceding simplified nodes. total represents the number of simplified nodes in the code vector graph. n1 is based on computing the reachability of a simplified node that is known to have data flow or control flow dependencies. In other words, the destructibility metric x n1 Is a value between 0 and 1, where 0 indicates if a simplification node is affected or broken by the attack in any way, it will not affect any part of the code. A simplification node with a value of 1 indicates that the code has a high chance of causing side effects / breakage in the code.
[0145] S412, multiply the n1th simplified node vector by the destructibility measure x of the n1th simplified node n1 Get channel value one.
[0146] S413, traverse N1 simplified nodes to obtain N1 channel values 1. Arrange the N1 channel values 1 in the order of code lines to obtain a destructibility measurement channel.
[0147] 2. The method for calculating the feature vector channel includes the following steps:
[0148] S421. Calculate the eigenvector centrality EC of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is:
[0149]
[0150] Where α represents the proportional constant. β represents the eigenvector centrality EC of the n1th simplified node n1 The initial value of A. A represents the adjacency matrix, and its eigenvalue is represented by λ, and α must satisfy the condition α<1 / λ. n1j1 Represents the edge weight or connection relationship between the n1th simplified node and the j1th simplified node. j1 Represents the eigenvector centrality EC of the j1th simplified node j1 , where the j1th simplified node and the n1th simplified node are adjacent nodes, j1∈[1, N1]. The eigenvector centrality EC n1 It means that the centrality of a simplified node depends on the centrality of its adjacent simplified nodes. In other words, the importance of a simplified node is proportional to the importance of the simplified nodes it is connected to.
[0151] S422, multiply the vector of the n1th simplified node by the eigenvector centrality EC of the n1th simplified node n1 Get channel value two.
[0152] S423, traverse N1 simplified nodes to obtain N1 channel values 2. Arrange the N1 channel values 2 in the order of code lines to obtain a feature vector channel.
[0153] 3. The method for calculating the approach channel includes the following steps:
[0154] S431, calculate the closeness centrality CC of the n1th simplified node n1 , and its calculation formula is:
[0155]
[0156] Where, d n1j2 represents the shortest distance between the n1th simplified node and the j2th node. N represents the number of all simplified nodes in the CSG. Closeness centrality CC n1 Indicates how close a node is to all other nodes in the code vector graph. It is calculated as the average of the shortest path lengths from the simplified node to every other simplified node in the code vector graph. The smaller the average shortest distance of a simplified node, the greater the closeness centrality of the simplified node.
[0157] S432, multiply the vector of the n1th simplified node by the closeness centrality CC of the n1th simplified node n1 Get channel value three.
[0158] S433, traverse N1 simplified nodes to obtain N1 channel values 3. Arrange the N1 channel values 3 in the order of code lines to obtain a close channel.
[0159] So far, three channels are obtained: destructibility metric channel, feature vector channel and proximity channel, and a weight image is generated based on the RGB three color channels. Specifically, an image usually has three color channels (i.e., red (R), green (G), and blue (B)), which work together to produce a complete image. The destructibility metric channel, feature vector channel and proximity channel correspond to the three color channels respectively, so that the destructibility metric x used in this method n1 , Eigenvector centrality EC n1 and closeness centrality CC n1 The importance of all lines of code in the software code is calculated from three different aspects in a corresponding manner, so that the contribution of different lines of code to the semantics of the software program can be fully considered. The final result is a weighted image with the importance of all lines of code.
[0160] S5. Adjust the weight image to the input specification requirements of the CNN model and perform a binary classification task to complete the vulnerability detection of the software code.
[0161] In this step, a CNN model with an attention mechanism is used to identify weight images, that is, to detect vulnerabilities. The CNN model has been trained by giving the generated images in advance. Since CNN uses images of the same size as input, and the number of lines of code for different functions is different, it is necessary to adjust the size of the input weight image. In this embodiment, 100 lines of code are used as the threshold for generating weight images. When the number of lines of code in a software program is less than 100, zero vectors are filled at the end of the code line (i.e., zero vectors are filled at the end of the destructibility metric channel, feature vector channel, and proximity channel) until the code line is filled to the threshold size. When the number of lines of code in a software program is greater than 100, the code lines at the end are deleted in turn (i.e., the end data of the destructibility metric channel, feature vector channel, and proximity channel are deleted). The weight image size of the final input CNN model is 3*100*768. Among them, "3" corresponds to three channels (i.e., destructibility metric channel, feature vector channel, and proximity channel), "100" corresponds to the threshold of the number of lines of code, and "768" represents the dimension of node vector embedding. The weighted image is calculated by the CNN model, and then convolution, pooling, and fully connected layer dimension conversion are performed. Finally, the softmax layer is applied to normalize the output probability and perform a binary classification task. The result of the binary classification task is yes or no, thereby obtaining the vulnerability detection result of the software code.
[0162] The code vulnerability detection method based on graph simplification (hereinafter referred to as GSVD) provided by the present invention has the advantages of high vulnerability detection efficiency and high detection accuracy. A comparative experiment was conducted for this purpose. GSVD was tested with some existing vulnerability detection methods. Existing vulnerability detection methods include VulDeePecker, SySeVR, a detection method based on the CodeBERT model, a detection method based on the GraphCodeBERT model, a detection method based on the ReVeal model, a detection method based on the ReGVD model, a detection method based on the Devign model, and a detection method based on the VulCNN model. The experimental results are shown in Table 4.
[0163] Table 4: Data results of comparison between GSVD and other methods
[0164]
[0165] As can be seen from Table 4, the accuracy (Acc), recall (R), and F1 value of GSVD are higher than those of existing vulnerability detection methods. Analysis: VulDeePecker and SySeVR extract code information by slicing, while CodeBERT and GraphCodeBERT regard the code as a sequence, both of which lack structural information about the program. Therefore, the extracted defect features are not sufficient, which affects the prediction of defects. In contrast, GSVD represents the program as a graph and retains the structural features in the program. Therefore, GSVD has higher accuracy, recall, and F1 value. In ReVeal and ReGVD, the edges are ordered, which is essentially still a natural sorting relationship rather than an inherent structural relationship of the code itself. Devign uses CPG (code property graph) to represent the code and uses the GGNN model to detect vulnerabilities, but many nodes and edges in CPG are irrelevant to the semantics of the vulnerability, which may reduce its accuracy and increase time overhead. GSVD uses PDG as the input of the model, simplifies PDG according to the type of node to obtain CSG, and uses a pre-trained model for vector embedding, which better enhances the representation of global information.
[0166] Then, the GSVD provided by the present invention is verified in terms of detection efficiency. GSVD is compared with some existing vulnerability detection methods, and the time required to detect vulnerabilities in 1000 software programs is uniformly calculated. The experimental results are as follows: Figure 3 As shown. Figure 3 It can be seen that GSVD only takes about 25 seconds to classify 1,000 images. Among some existing vulnerability detection methods, VulDeePecker uses BiLSTM to learn vulnerability features, which is a variant of recurrent neural network (RNN) that specializes in processing sequential data and therefore has a relatively small time overhead. Although Reveal only uses MLP to classify vulnerabilities, it still requires some time overhead due to the introduction of the smoke algorithm. Devign combines different code representations into a composite graph, which has a relatively high time overhead in vulnerability detection because the composite graph contains a large number of edges. DeepKoko selects the GCN model, which requires the execution of additional convolutional layers compared to the GNN model, so it has a larger time overhead. Therefore, the comparison results show that the GSVD provided by the present invention has a competitive advantage in vulnerability detection efficiency, and further proves the effectiveness and practicality of GSVD in vulnerability detection tasks, providing a more effective and feasible solution strategy for vulnerability detection in large-scale projects.
[0167] In some other embodiments, a code vulnerability detection system based on graph simplification is also proposed, which uses the code vulnerability detection method based on graph simplification as described above. The code vulnerability detection system based on graph simplification includes a preprocessing module, a graph simplification module, a node embedding module, an image generation module and a classification module.
[0168] Specifically, the preprocessing module is used to preprocess the original software code code0 to remove noise, and obtain the preprocessed software code code1. It is also used to perform image generation processing on the preprocessed software code to obtain the program dependency graph PDG.
[0169] The graph simplification module is used to simplify the PDG to obtain the program simplified graph CSG.
[0170] The node embedding module is used to perform vector embedding on the N1 simplified nodes in the CSG, obtain N1 simplified node vectors, and update the CSG into a code vector graph.
[0171] The image generation module is used to perform centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image.
[0172] The classification module is used to adjust the weight image to the input specification requirements of the CNN model and perform binary classification tasks.
[0173] In some other embodiments, an electronic device is also provided. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the code vulnerability detection method based on graph simplification are implemented.
[0174] Some other embodiments also provide a computer-readable storage medium that stores a computer program that implements the steps of the above-mentioned code vulnerability detection method based on graph simplification when the computer program is executed by a processor.
[0175] Some other embodiments also provide a software program product, which includes program instructions, and when the software program product is run on an electronic device, the electronic device executes the steps of the code vulnerability detection method based on graph simplification as described above.
[0176] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0177] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.< / op> < / operator>
Claims
1. A code vulnerability detection method based on graph simplification, characterized in that: It includes the following steps: S1, preprocessing the original software code code0 to remove noise, and obtaining the preprocessed software code code1; wherein code1 includes K1 lines of code; Perform image generation processing on code1 to obtain the program dependency graph PDG; Among them, PDG includes K1 original nodes and K2 original node edges; The k1th original node represents the k1th line of code in code1; k1∈[1, K1]; K2 original node edges represent the dependency relationship between K1 original nodes; S2, perform graph simplification on PDG to obtain a program simplified graph CSG; CSG includes N1 simplified nodes and N2 simplified node edges; N1<K1, N2<K2; The graph simplification processing method includes: S21, merging adjacent original nodes in the PDG according to the node type, the merged simplified nodes inherit the original node edges of the original nodes before the merger, and form simplified node edges; S22, remove the noise nodes and free nodes in the PDG, and delete the original node edges connected to the noise nodes and free nodes; S3, embed the N1 simplified nodes in the CSG into vectors to obtain N1 simplified node vectors, and update the CSG into a code vector graph; S4, performing centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image; S5 adjusts the weight image to the input specification requirements of the CNN model and performs a binary classification task to complete the vulnerability detection of the software code.
2. The code vulnerability detection method based on graph simplification according to claim 1 is characterized in that: In S21, the method for merging adjacent original nodes according to node type includes the following steps: Determine whether the adjacent original nodes meet the merging rules; If the adjacent original nodes meet the merging rules, the adjacent original nodes are merged; The merging rules include: If the parent node type is an expression statement and the child node type is also an expression statement, they will be merged; If the parent node type is an identifier declaration statement and the child node type is an identifier, they are merged; If the parent node type is a conditional statement, all child nodes of any type will be merged; If the parent node type is a for loop statement and the child node type is a parameter list, they will be merged; If the parent node type is a function call statement and the child node type is an identifier or parameter list, they are merged; The method of merging adjacent original nodes includes retaining the original node of the parent node type, deleting the original node of the child node type, and retaining the original node edge corresponding to the original node of the child node type; Among the adjacent original nodes, the parent node is the original node located in front of the edge of the original node, and the child node is the original node located behind the edge of the original node.
3. The code vulnerability detection method based on graph simplification according to claim 1 is characterized in that: In S22, the method for determining the noise node includes the following steps: S221, determining the attributes of the predefined noise node; S222, comparing the attributes of the K1 original nodes and the predefined noise nodes in the PDG one by one; If the attribute of an original node is consistent with the attribute of a predefined noise node, the original node is determined to be a noise node.
4. The code vulnerability detection method based on graph simplification according to claim 3 is characterized in that: The method for determining the predefined noise node comprises the following steps: S2211, preprocessing the sample software code to remove noise, and performing image generation processing on the preprocessed sample software code to obtain a program dependency graph PDG'; Among them, the sample software code is software code that is known to have vulnerabilities; S2212, counting the node types in PDG'; S2213, conduct an ablation experiment on each node type of PDG', and regard the node types that have a negative impact on the detection results as suspicious node types; S2214, performing an ablation experiment on each node in the suspicious node type, and determining the nodes that have a negative impact on the detection result as noise nodes; The properties of the noise node are recorded and saved as a predefined noise node.
5. The code vulnerability detection method based on graph simplification according to claim 1 is characterized in that: In S22, the method for determining a free node includes the following steps: S2215, calculating the degrees of K1 original nodes in the PDG; S2216, the original node with a degree of 0 is regarded as a free node.
6. The code vulnerability detection method based on graph simplification according to claim 1 is characterized in that: In S4, the method for generating a weight image comprises the following steps: S41, calculate the destructibility metrics x1~x1 of N1 simplified nodes in the code vector graph N1 , eigenvector centrality EC1~EC N1 and closeness centrality CC1~CC N1 ; S42, based on the destructibility metrics x1~x1 of N1 simplified nodes N1 The destructibility measurement channel is calculated; According to the eigenvector centrality of N1 simplified nodes EC1~EC N1 Calculate the characteristic vector channel; According to the closeness centrality of N1 simplified nodes CC1~CC N1 The approach channel is calculated; S43, generating a weight image according to the destructibility metric channel, the feature vector channel, and the proximity channel; The weight image is used to characterize the importance of the N1 row simplification nodes.
7. The code vulnerability detection method based on graph simplification according to claim 6 is characterized in that: The method for calculating the destructibility metric channel includes the following steps: S411, calculate the destructibility measure x of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is: Where n reachable Indicates the number of simplified node edges between the n1th simplified node and its adjacent simplified nodes, n total Represents the number of all simplified nodes in the code vector graph; S412, multiplying the n1th simplified node vector by the destructibility measure x of the n1th simplified node n1 Get channel value one; S413, traverse N1 simplified nodes to obtain N1 channel values 1; Arrange the N1 channel values in the order of code lines to obtain the destructibility measurement channel.
8. The code vulnerability detection method based on graph simplification according to claim 6 is characterized in that: The method for calculating the feature vector channel includes the following steps: S421, calculate the eigenvector centrality EC of the n1th simplified node n1 , n1∈[1,N1], the calculation formula is: Where α represents the proportional constant, and β represents the eigenvector centrality EC of the n1th simplified node. n1 The initial value of; A represents the adjacency matrix; A n1j1 represents the edge weight or connection relationship between the n1th simplified node and the j1th simplified node; x j1 Represents the eigenvector centrality EC of the j1th simplified node j1 , the j1th simplified node and the n1th simplified node are adjacent nodes, j1∈[1,N1]; S422, multiply the vector of the n1th simplified node by the eigenvector centrality EC of the n1th simplified node n1 Get channel value 2; S423, traverse N1 simplified nodes to obtain N1 channel values 2; Arrange the N1 channel values in the order of code lines to obtain the feature vector channel; The method for calculating the approach channel includes the following steps: S431, calculate the closeness centrality CC of the n1th simplified node n1 , and its calculation formula is: Where, d n1j2 represents the shortest distance between the n1th simplified node and the j2th node; j2∈[1, N1]; N represents the number of all simplified nodes in CSG; S432, multiply the vector of the n1th simplified node by the closeness centrality CC of the n1th simplified node n1 Get channel value three; S433, traverse N1 simplified nodes to obtain N1 channel values three; Arrange the N1 channel values in the order of code lines to obtain the approximate channel.
9. The code vulnerability detection method based on graph simplification according to claim 1 is characterized in that: In S5, the method of adjusting the weight image to the CNN model input specification requirement includes the following steps: If the size of the weight image is less than the threshold, zero vectors are filled at the end of the destructibility metric channel, feature vector channel, and proximity channel; If the size of the weight image is larger than the threshold, the destructibility metric channel, the feature vector channel, and the end data of the near channel are deleted.
10. A code vulnerability detection system based on graph simplification, characterized in that: It uses the code vulnerability detection method based on graph simplification as described in any one of claims 1 to 9; The code vulnerability detection system based on graph simplification includes: A preprocessing module is used to preprocess the original software code code0 to remove noise, and obtain the preprocessed software code code1; and is also used to perform image generation processing on the preprocessed software code to obtain a program dependency graph PDG; A graph simplification module, which is used to simplify the PDG to obtain a program simplified graph CSG; A node embedding module is used to embed vectors of N1 simplified nodes in CSG, obtain N1 simplified node vectors, and update CSG into a code vector graph; An image generation module, which is used to perform centrality analysis on the code vector graph based on simplified nodes, simplified node edges and simplified node vectors to generate a weight image; The classification module is used to adjust the weight image to the input specification requirements of the CNN model and perform a binary classification task.
Citation Information
Patent Citations
Similarity detection method for unknown vulnerability discovery based on patch information
CN108268777A
Intrusion detection method based on space-time characteristics and attention mechanism
CN114697096A
Intelligent contract vulnerability detection method based on self-supervised learning
CN116340951A
Industrial protocol software source code vulnerability detection and positioning method based on association diagram
CN117744085A
Source code vulnerability detection method and system based on code edge
CN117932609A