Source code program dependency graph pruning optimization method for vulnerability detection
By filtering key nodes in the program dependency graph and generating dependency subgraphs, and combining with the graph neural network model for vulnerability detection, the problem of huge scale and large redundant information in the existing technology is solved, and efficient and accurate vulnerability detection effect is achieved.
Patent Information
- Application Number
- CN202510230257.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the large scale of the program dependency graph and a lot of redundant information are caused by high computational complexity, low analysis efficiency, and difficulty in achieving efficient and accurate vulnerability detection in large-scale code bases.
A source code program dependency graph pruning optimization method for vulnerability detection is proposed. By filtering key nodes such as pointer operations, array access, sensitive API calls and integer overflow expressions, the source code dependency subgraph is generated, and vulnerability detection is carried out in combination with deep learning graph neural network model.
It significantly reduces irrelevant nodes and edges in the dependency graph, improves analysis efficiency and detection accuracy, effectively reduces vulnerability false alarm rate, and is suitable for fast vulnerability detection in large-scale code bases.
Smart Images

Figure CN120197176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a source code program dependency graph pruning optimization method for vulnerability detection. Background Art
[0002] With the rapid expansion of software development and the continuous increase in complexity, software vulnerabilities are becoming more and more serious threats to information security and have become one of the key risk factors affecting system stability and data security. The existence of vulnerabilities may not only lead to sensitive data leakage and system function interruption, but may also be maliciously exploited by attackers, leading to serious network attacks. Therefore, how to detect vulnerabilities efficiently and accurately has become an important issue that needs to be solved in the field of computer security.
[0003] Software vulnerabilities usually refer to defects or errors in the design, coding or operation of source code. By detecting vulnerabilities in source code, potential risks can be discovered during the development phase, thereby reducing the possibility of vulnerabilities being maliciously exploited. Existing source code vulnerability detection technologies are mainly divided into dynamic analysis and static analysis. Dynamic analysis detects vulnerabilities through the program running status and can accurately capture vulnerability features, but it relies on test case coverage and the configuration of the running environment, which limits its applicability in complex systems and large-scale code bases. In contrast, static analysis does not need to rely on the program running environment and directly mines vulnerabilities through source code, so it is suitable for comprehensive detection of large-scale codes.
[0004] One of the core technologies of static analysis is the generation of Program Dependence Graph (PDG). The PDG provides a comprehensive structured representation for vulnerability detection by describing the control flow and data flow relationships in the code. However, the size and complexity of the dependency graph usually expands rapidly with the increase of code volume, which becomes one of the key technical challenges faced by static analysis methods. Summary of the invention
[0005] To address the computational complexity issue caused by the scale expansion of the program dependence graph, researchers have proposed various optimization methods, among which pruning optimization is an effective technique. Pruning optimization reduces the scale and complexity of the dependence graph by streamlining redundant nodes and edges in the program dependence graph and retaining only the key information directly related to vulnerability detection, thus significantly improving the analysis efficiency and detection accuracy. Pruning optimization not only simplifies the structure of the program dependence graph but also lays a foundation for the application of subsequent deep learning techniques. Based on the pruned dependence graph, researchers combine deep learning techniques, especially the Graph Neural Network (GNN). GNN extracts deep features from the structure and semantic information of the program dependence graph and learns the complex dependence relationships between nodes through neighborhood aggregation, thereby capturing features related to vulnerabilities.
[0006] The present invention proposes a pruning optimization method for the source code extended program dependence graph for vulnerability detection. Aiming at the problems of large scale and a large amount of redundant information in the program dependence graph in the prior art, an efficient pruning optimization strategy is designed. By screening nodes strongly related to vulnerability detection, such as pointer operations, array accesses, sensitive API calls, and integer overflow expressions, and retaining only the core features, the irrelevant nodes and edges in the dependence graph are significantly reduced. On this basis, combined with the graph neural network model of deep learning, efficient and accurate vulnerability detection is realized.
[0007] To achieve the above object, the present invention provides the following technical solution: A pruning method for the source code program dependence graph for vulnerability detection, comprising the following steps:
[0008] Step 1, download the source code data set, denote the source code data set as data set A, perform attribute screening on the data set A to obtain the filtered data set, and denote the filtered data set as data set B;
[0009] Step 2, preprocess the data set B to remove the irrelevant information in the data set B to form sample one;
[0010] Step 3, convert sample one into a program dependence graph representation to form sample two; sample two can comprehensively present the semantic logic, syntax features, and program structure characteristics of the code of sample two through the combined data structure of control dependence, data dependence, and abstract syntax tree;
[0011] Step 4, perform pruning on sample two based on pointer nodes, array nodes, sensitive API call nodes, and integer overflow expression nodes to generate a program dependence subgraph denoted as sample three;
[0012] Step 5: Based on the Word2Vec model of word vectors, perform vectorization training on the code semantic units in Sample 3 to generate a vector representation of the code snippet that fuses grammatical and semantic features, denoted as Sample 4;
[0013] Step 6: Use a graph neural network to train and predict Sample 4 to obtain the vulnerability detection result of the source code.
[0014] Preferably, Step 1 is specifically implemented according to the following steps:
[0015] Step 1.1: Download the source code dataset file containing multiple attributes, denoted as Dataset A;
[0016] Step 1.2: Screen out the data with the value of the attribute vul field being 1 from Dataset A to construct the vulnerability dataset B; Concatenate the project name project and the commit version number commit_id in Dataset B to generate a unique file identifier in the format of {project}_{commit_id};
[0017] Step 1.3: Store the vulnerable code of each sample in Dataset B as a vulnerability file, and store the repaired code of each sample in Dataset B as a repair file;
[0018] Step 1.4: By comparing the vulnerability file and the repair file line by line, extract the line numbers of the differences between the compared vulnerability file and the repair file, and generate an annotation file with the line numbers of the differences.
[0019] Preferably, Step 1.4 is specifically implemented according to the following steps:
[0020] Step 1.4.1: Use a comparison tool to compare the vulnerability file and the repair file line by line to generate a difference list; Among them, if a line starts with the "-" symbol, it means a deleted line that only exists in the vulnerability file, and if a line starts with the "+" symbol, it means an added line that only exists in the repair file;
[0021] Step 1.4.2: In the difference list, if a line is a deleted line starting with the "-" symbol, record the line number of this line in the annotation file.
[0022] Preferably, Step 2 is specifically implemented according to the following steps:
[0023] Step 2.1: The preprocessing of Dataset B is as follows: Match the comments in the target file and the repair file through regular expressions to obtain the processed source code, and the comments include: single-line comment symbols and multi-line comment blocks;
[0024] Step 2.2: Create a set of keyword collections, including keywords in C language, preprocessor directives, built-in functions, and API names;
[0025] Step 2.3: Match the processed source code through a lexical analyzer and regular expression rules to obtain function names and variable names;
[0026] Step 2.4: For function names not in the keyword set, generate new symbols in the order they appear in the code and replace the original function names;
[0027] Step 2.5: If the variable name is not in the keyword set and is not the number of command-line arguments or the command-line argument vector, generate new symbols for the variable names not in the keyword set in the declaration order to obtain Sample One.
[0028] Preferably, the specific implementation of Step 3 is as follows:
[0029] Step 3.1: Parse Sample One into a binary file and convert the binary file into files in DOT and JSON formats;
[0030] Step 3.2: Extract all edges from the DOT file, extract all nodes from the JSON file, and filter out the nodes that do not contain the line number attribute among the nodes to form filtered nodes;
[0031] Step 3.3: Sort the filtered nodes according to the column numbers and line numbers of the code in the program to form sorted nodes; through the abstract syntax tree, data dependence graph, and control dependence graph, match and integrate the sorted nodes with all edges to construct a program dependence graph, denoted as Sample Two.
[0032] Preferably, the specific implementation of Step 4 is as follows:
[0033] Step 4.1: Obtain pointer nodes, array nodes, sensitive API call nodes, and integer overflow expression nodes from Sample Two to obtain the focus points of the code slice;
[0034] Step 4.2: After obtaining the focus points, generate a list of program slices that is consistent with the actual order of the code by traversing the parent nodes in the abstract syntax tree and combining the forward and backward slicing methods;
[0035] Step 4.3: Traverse each subgraph in the list of program slices for vulnerability checking to obtain the checked subgraph; perform vulnerability checking on the annotation file in Step 1.4 above;
[0036] Step 4.4: For the checked subgraph in Step 4.3 above, initialize a directed graph object and create corresponding nodes in the graph for each node in the checked subgraph; then, traverse the edge information of each node, record the type of the edge and the connected nodes, and add them to the directed graph to obtain the processed directed graph;
[0037] Step 4.5: Check the edges in the processed directed graph to ensure that the source node and target node of each edge exist, and generate the final sub-graph, denoted as Sample Three.
[0038] Preferably, Step 4.2 is specifically implemented according to the following steps:
[0039] Step 4.2.1: For each code concern point, perform the following structured traversal operations. First, obtain the direct parent node of this concern point in the abstract syntax tree. If the current parent node has a data dependency edge or is a method definition node, directly add it to the concern point list. Otherwise, traverse upward along the abstract syntax tree to the first ancestor node that meets the above conditions. Perform deduplication detection on the ancestor node. If it does not exist in the concern point list, dynamically append it to form a new concern point list.
[0040] Step 4.2.2: Process all nodes in the new concern point list in a loop. For each node in the new concern point list, perform slicing by combining backward and forward methods to obtain the successor and predecessor nodes of this node, and add the successor and predecessor nodes to the slicing list for integration.
[0041] Step 4.2.3: When integrating the slicing list, exclude unimportant or incomplete successor and predecessor nodes, and sort all successor and predecessor nodes according to the line numbers to generate a program slicing list that is consistent with the actual code order.
[0042] Preferably, Step 5 is specifically implemented according to the following steps:
[0043] Step 5.1: Traverse the node information of Sample Three sub-graph, and split the code string of each node of Sample Three sub-graph into lexical tokens.
[0044] Step 5.2: Train the lexical tokens according to the Word2Vec model to generate vector representations of the tokens.
[0045] Step 5.3: Convert Sample Three into a graph structure to generate feature vectors of all nodes of Sample Three sub-graph.
[0046] Step 5.4: Traverse the edge information of Sample Three to obtain the vector features of each edge.
[0047] Step 5.5: Save the vector representations of the tokens, the feature vectors of all nodes of Sample Three sub-graph, and the vector features of each edge to Sample Four.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] (1) Based on the traditional program dependence graph, by refining the node granularity and combining lexical and syntactic analysis, the source code is decomposed into tokens (such as keywords, variable names, operators, etc.), making up for the deficiency of the traditional program dependence graph in capturing syntactic information.
[0050] (2) During the pruning process, around four types of benchmark points: pointers, arrays, sensitive API calls, and integer overflow expressions, the program dependence graph is pruned and optimized to generate the dependence structure of the source code, solving the problems of large computational overhead and insufficient detection accuracy of traditional methods when dealing with large-scale code files.
[0051] (3) Refine the granularity of source code vulnerability detection to the statement level, solving the problem that further positioning is still required after discovering vulnerabilities at the function level in previous studies.
[0052] Train and predict the generated dependence subgraph through a graph neural network, thereby effectively reducing the false positive rate of vulnerabilities and improving the efficiency and accuracy of vulnerability detection. Description of the Drawings
[0053] Figure 1 It is the processing flow chart in the source code program dependence graph pruning and optimization method for vulnerability detection of the present invention. Detailed Embodiments
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0055] The present invention provides a source code program dependence graph pruning and optimization method for vulnerability detection, as Figure 1 shown, and can be specifically implemented according to the following steps:
[0056] Step 1, download the source code data set, denote the source code data set as data set A, and perform attribute screening on the data set A to obtain the filtered data set, and denote the filtered data set as data set B;
[0057] Step 2, preprocess the data set B to remove the irrelevant information in the data set B to form sample one;
[0058] Step 3, convert the sample one into a program dependence graph representation to form sample two; the sample two can comprehensively present the semantic logic, syntactic features, and program structure characteristics of the code of the sample two through a joint data structure of control dependence, data dependence, and abstract syntax tree.
[0059] Step 4, perform pruning based on pointer nodes, array nodes, sensitive API call nodes, and integer overflow expression nodes in Sample 2, and generate a program dependence subgraph denoted as Sample 3;
[0060] Step 5, based on the Word2Vec model of word vectors, perform vectorization training on the code semantic units in Sample 3 to generate a vector representation of code snippets that integrates syntactic and semantic features, denoted as Sample 4;
[0061] Step 6, use a graph neural network to train and predict Sample 4 to obtain the vulnerability detection result of the source code.
[0062] For the above, the specific implementation of Step 1 is as follows:
[0063] Step 1.1, download the source code dataset CSV file containing multiple attributes, denoted as Dataset A;
[0064] Step 1.2, screen out the data with the attribute vul field value of 1 from Dataset A to construct a vulnerability dataset B; concatenate the project name project and the commit version number commit_id in Dataset B to generate a unique file identifier in the format of {project}_{commit_id}, such as Android_0b23c8;
[0065] Step 1.3, store the vulnerable code func_before of each sample in Dataset B as a vulnerability file F_vul, and store the repaired code func_after of each sample in Dataset B as a repaired file F_novul; the suffixes of the vulnerability file F_vul and the repaired file F_novul are unified as.c; the file name generation rule is as follows: for the vulnerability file, use 1_ as the prefix, such as the file 1_Android_0b23c8.c, denoted as F_vul; for the repaired file, use 0_ as the prefix, the file 0_Android_0b23c8.c, denoted as F_novul;
[0066] Step 1.4, by comparing the vulnerability file F_vul and the repaired file F_novul line by line, extract the line numbers of the differences between the vulnerability file F_vul and the repaired file F_novul, and generate an annotation file with the line numbers of the differences.
[0067] For the above, the specific implementation of Step 1.4 is as follows:
[0068] Step 1.4.1: Use a comparison tool (such as the difflib library or Unix diff) to compare the vulnerability file F_vul and the fixed file F_novul line by line, and generate a unified-format difference list diff (such as: @@-15,6+16,7@@). Among them, the symbol "@@" marks the start of a code difference block; lines starting with the symbol "-" indicate deleted lines that only exist in F_vul; lines starting with the symbol "+" indicate added lines that only exist in F_novul; if no symbol is marked, it represents the common retained code segment of the two files.
[0069] Step 1.4.2: In the difference list diff, determine the starting position of the difference block by looking for lines starting with "@@"; in each difference block, if a line starts with the "-" symbol (deleted line), record the line number of that line in the annotation file diff_labels.
[0070] The above-mentioned Step 2 is specifically implemented according to the following steps:
[0071] Step 2.1: The preprocessing of the dataset B is as follows: Use regular expressions to match the comments in the target file F_vul and the fixed file F_novul to obtain the processed source code. The comments include single-line comment symbols (such as / / ) and multi-line comment blocks (such as / *...* / ); when it is detected that the comment content does not contain a newline character, perform the operation of clearing the comment content and retain the original newline character (such as \n); when it is detected that the comment content contains N newline characters, replace the comment block with N consecutive newline characters to maintain the integrity of the line-level topology of the source code; achieve the elimination of the interference of comment content on subsequent code feature analysis while maintaining the integrity of the code line number mapping relationship.
[0072] Step 2.2: Create a set of keyword terms, including 44 keywords in the C language, preprocessor directives, built-in functions, and API names.
[0073] Step 2.3: Match the processed source code through a lexical analyzer and regular expression rules to obtain function names and variable names; the difference between a function and a variable is that there are parentheses after a function.
[0074] Step 2.4: For function names that are not in the set of keyword terms, generate new symbols (such as FUN_1) in the order they appear in the code and replace the original function names.
[0075] Step 2.5, if the variable name is not in the set of keyword terms and is not the number of command-line arguments "argc" or the command-line argument vector "argv", then generate a new symbol (such as VAR_1) for the variable name that is not in the set of keyword terms in the order of declaration. After completing all preprocessing, obtain Sample One.
[0076] As described above, Step 3 is specifically implemented according to the following steps:
[0077] Step 3.1, parse Sample One into a binary bin file, and convert the binary bin file into files in DOT and JSON formats; where, the nodes of the DOT file represent each line of code in the source code, and the edges represent data dependencies and control dependencies; the JSON file represents the abstract syntax tree of the code, including: the program structure and syntax information of the source code;
[0078] Step 3.2, extract all edges from the DOT file, extract all nodes from the JSON file, and filter out the nodes that do not contain the line number attribute in the nodes to form filtered nodes;
[0079] Step 3.3, sort the filtered nodes according to the column numbers and line numbers of the code in the program to form sorted nodes; through the abstract syntax tree, data dependency graph, and control dependency graph, match and integrate the sorted nodes with all edges to construct a program dependency graph, denoted as Sample Two.
[0080] As described above, Step 4 is specifically implemented according to the following steps:
[0081] Step 4.1, obtain pointer nodes, array nodes, sensitive API call nodes, and integer overflow expression nodes from Sample Two to obtain the focus points of the code slice;
[0082] Step 4.2, after obtaining the focus points, generate a list of program slices that is consistent with the actual order of the code by traversing the parent nodes in the abstract syntax tree and combining the forward and backward slicing methods;
[0083] Step 4.3, traverse each subgraph in the list of program slices for vulnerability checking to obtain the subgraphs after checking; perform vulnerability checking on the annotation file in Step 1.4 above, and check each vulnerable slice file with the prefix "1_" one by one. If there is no vulnerable line in the file, update the file name prefix to "0_";
[0084] Step 4.4, for the subgraphs after checking in Step 4.3 above, initialize a directed graph object, and create corresponding nodes in the graph for each node in the subgraphs after checking; then, traverse the edge information of each node, record the type of the edge and the connected nodes, and add them to the directed graph to obtain the processed directed graph;
[0085] Step 4.5: Verify the edges in the processed directed graph to ensure that the source node and target node of each edge exist, generate the final subgraph and save it in DOT format, and finally obtain the subgraph with the complete structure, denoted as Sample Three.
[0086] Specifically, the above Step 4.2 is implemented according to the following steps:
[0087] Step 4.2.1: For each code concern point, perform the following structured traversal operations: First, obtain the direct parent node of this concern point in the abstract syntax tree; if the current parent node has a data dependency edge or is a method definition node, directly add it to the concern point list; otherwise, traverse upward in the abstract syntax tree to the first ancestor node that meets the above conditions; perform duplicate detection on the ancestor node, and if it does not exist in the concern point list, dynamically append it to form a new concern point list.
[0088] Step 4.2.2: Process all nodes in the new concern point list in a loop; for each node in the new concern point list, perform slicing in a combined way of backward and forward to obtain the successor and predecessor nodes of this node, and add the successor and predecessor nodes to the slicing list for integration.
[0089] Step 4.2.3: When integrating the slicing list, exclude unimportant or incomplete successor and predecessor nodes, ensure that each node appears only once, and sort all successor and predecessor nodes according to the line numbers to make the slicing result consistent with the actual code order, and generate a program slicing list consistent with the actual code order.
[0090] Specifically, the above Step 5 is implemented according to the following steps:
[0091] Step 5.1: Traverse the node information of Sample Three subgraph, and split the code string of each node in Sample Three subgraph into lexical tokens (Tokens), such as operators, variable names, keywords, etc.
[0092] Step 5.2: Train the lexical tokens (Tokens) according to the Word2Vec model to generate the vector representation of the tokens.
[0093] Step 5.3: Convert Sample Three into a graph structure and generate the feature vectors of all nodes in Sample Three subgraph.
[0094] Step 5.4: Traverse the edge information of Sample Three to obtain the vector features of each edge.
[0095] Step 5.5: Save the vector representation of the tokens, the feature vectors of all nodes in Sample Three subgraph, and the vector features of each edge as Sample Four; Sample Four is a JSON format file.
[0096] The above-mentioned step 6 is specifically implemented according to the following steps:
[0097] Step 6.1: Randomly divide the JSON format file of sample four into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0098] Step 6.2: Use a graph neural network to train the data set composed of the training set and the validation set, and record the best weights of the graph neural network model;
[0099] Step 6.3: During the evaluation process, load the best weights pre-trained by the model, and then perform performance evaluation of vulnerability detection on the test set; The evaluation metrics include accuracy, precision, recall, and F1 score, and these metrics are recorded.
[0100] The above-mentioned step 6.2 is specifically implemented according to the following steps:
[0101] Step 6.2.1: Read and parse the data of the training set and the validation set, divide the data into multiple batches according to the set batch size, and convert it into a batch graph object suitable for processing by the graph neural network model;
[0102] Step 6.2.2: When initializing the graph neural network model, it receives the input feature dimension, output feature dimension, number of message passing steps, and maximum number of edge types as parameters. Through these parameters, the model initializes the gated graph convolutional layer and the fully connected layer;
[0103] Step 6.2.3: During the training process, the initialized model performs forward propagation, extracts the graph structure, node features, and edge types from the input batch graph object, and moves this data to the GPU for accelerated calculation;
[0104] Step 6.2.4: The graph neural network model performs message passing through the gated graph convolutional layer to update the features of each node. Then, the features of all nodes of each graph are summed to obtain the global feature representation of the graph. The global feature is mapped to a scalar value (representing the classification score) through the fully connected layer, and finally the sigmoid function is used to compress the output into the range of [0,1] for the final result output of the binary classification task (for example, determining whether there is a vulnerability);
[0105] Step 6.2.5: After completing the forward propagation, calculate the loss value and perform backpropagation to obtain the gradient. Use an optimizer (such as Adam) to update the model parameters. Every certain number of steps (such as 50 steps), the system will enter the validation phase. During the validation phase, calculate the loss, F1 score, and accuracy on the validation set, and record these metrics for subsequent analysis;
[0106] Step 6.2.6, if the F1 score on the validation set is greater than 50 and the accuracy is higher than the historical best value, save the current model weights as the best model. When the performance on the validation set does not improve, the patience counter is incremented. Once the patience counter reaches the pre-set maximum value, the early stopping mechanism is triggered to terminate the training early to avoid overfitting.
[0107] The present invention provides a specific embodiment as follows:
[0108] Step 1.1, download the source code dataset CSV file containing 21 attributes, denoted as dataset A;
[0109] Step 1.2, screen out the data with the attribute vul field value of 1 from dataset A to construct the vulnerability dataset B; concatenate the project (project name) and commit_id (commit version number) in dataset B to generate a unique file identifier in the format of {project}_{commit_id}, such as Android_0b23c8;
[0110] Step 1.3, store the func_before (vulnerable code) and func_after (fixed code) of each sample in dataset B as independent files respectively, with the file suffix unified as.c. The file name generation rule is as follows: for the vulnerability file, use 1_ as the prefix, such as the file 1_Android_0b23c8.c, denoted as F_vul; for the fixed file, use 0_ as the prefix, the file 0_Android_0b23c8.c, denoted as F_novul;
[0111] Step 1.4, by comparing the vulnerability file F_vul and the fixed file F_novul line by line, extract the line numbers of the differences and generate an annotation file.
[0112] Step 1.4.1, use a comparison tool (such as the difflib library or Unix diff) to compare the vulnerability file F_vul and the fixed file F_novul line by line to generate a difference list diff with a unified format (such as: @@-15,6+16,7@@). Among them, the symbol "@@" marks the start of a code difference block; starting with the symbol "-" indicates a deleted line that only exists in F_vul; starting with the symbol "+" indicates an added line that only exists in F_novul, and if no symbol is marked, it indicates a common retained code segment of the two files;
[0113] Step 1.4.2, from the difference list diff, determine the starting position of the difference block by looking for the line starting with "@@"; in each difference block, if a line starts with the "-" symbol, record the line number of this line into the annotation file diff_labels;
[0114] Step 2.1, in the preprocessing stage of the dataset B, use regular expressions to match single-line comment symbols (such as / / ) and multi-line comment blocks (such as / *...* / ) in the target files F_vul and F_novul; when it is detected that the comment content does not contain line breaks, perform the operation of clearing the comment content and retain the original line breaks (such as \n); when it is detected that the comment content contains N line breaks, replace the comment block with N consecutive line breaks to maintain the integrity of the line-level topology of the source code; the technical effect of this step is to eliminate the interference of comment text on subsequent code feature analysis while maintaining the integrity of the code line number mapping relationship.
[0115] Step 2.2, create a set of keyword terms, including keywords in C language (44), preprocessor directives, built-in functions, and API names;
[0116] Step 2.3, match the processed source code through a lexical analyzer and regular expression rules to obtain function names and variable names; among them, the difference between a function and a variable is that there are parentheses after a function;
[0117] Step 2.4, for function names not in the set of keyword terms, generate new symbols (such as FUN_1) in the order of appearance in the code and replace the original function names;
[0118] Step 2.5, if the variable name is not in the set of keyword terms and is not the number of command-line arguments "argc" or the command-line argument vector "argv", then generate new symbols (such as VAR_1) for the variable names not in the set of keyword terms in the order of declaration. After all the preprocessing, obtain Sample One.
[0119] Step 3.1, parse Sample One into a binary bin file and convert it into files in DOT and JSON formats; among them, the nodes in the DOT file represent each line of code in the source code, and the edges represent data dependencies and control dependencies; the JSON file represents the abstract syntax tree of the code, including the program structure and syntax information of the source code;
[0120] Step 3.2, extract all edges from the DOT file, extract all nodes from the JSON file, and filter out the nodes that do not contain the line number attribute among the nodes to form filtered nodes;
[0121] Step 3.3, sort the filtered nodes according to the column numbers and line numbers of the code in the program to form sorted nodes; through the abstract syntax tree, data dependency graph, and control dependency graph, match and integrate the sorted nodes with all edges to construct a program dependency graph, denoted as Sample Two.
[0122] Step 4.1, obtain the pointer node, array node, sensitive API call node, and integer overflow expression node from Sample 2, and obtain the focus of the code slice;
[0123] Step 4.2.1, for each code focus, perform the following structured traversal operations: First, obtain the direct parent node of the focus in the abstract syntax tree; if the current parent node has a data dependency edge or is a method definition node, directly add it to the focus list; otherwise, traverse upward in the abstract syntax tree to the first ancestor node that meets the above conditions; perform deduplication detection on the ancestor node, and if it does not exist in the focus list, dynamically append it to form a new focus list;
[0124] Step 4.2.2, process all nodes in the new focus list in a loop; for each node in the new focus list, perform slicing in a combined way of backward and forward to obtain the successor and predecessor nodes of the node, and add the successor and predecessor nodes to the slice list for integration;
[0125] Step 4.2.3, when integrating the slice list, exclude unimportant or incomplete successor and predecessor nodes, ensure that each node appears only once, and sort all successor and predecessor nodes according to the line number to make the slice result consistent with the actual code order, and generate a program slice list consistent with the actual code order.
[0126] Step 4.3, traverse each subgraph in the program slice list for vulnerability checking to obtain the checked subgraph; perform vulnerability checking on the annotation file in the above step 1.4, and check each vulnerable slice file with the prefix name "1_" one by one. If there is no vulnerable line in the file, update the file name prefix to "0_";
[0127] Step 4.4, for the checked subgraph in the above step 4.3, initialize a directed graph object, and create corresponding nodes in the graph for each node in the checked subgraph; then, traverse the edge information of each node, record the type of the edge and the connected nodes, and add them to the directed graph to obtain the processed directed graph;
[0128] Step 4.5, verify the edges in the processed directed graph to ensure that both the source node and the target node of each edge exist, generate the final subgraph and save it in DOT format, and finally obtain the subgraph with a complete structure, denoted as Sample 3;
[0129] Step 5.1, traverse the node information of the Sample 3 subgraph, and split the code string of each node in the Sample 3 subgraph into lexical tokens (Tokens), such as operators, variable names, keywords, etc.;
[0130] Step 5.2, train the lexical tokens according to the Word2Vec model to generate vector representations of the tokens;
[0131] Step 5.3, convert Sample Three into a graph structure and generate feature vectors of the nodes of all subgraphs of Sample Three;
[0132] Step 5.4, traverse the edge information of Sample Three to obtain the vector features of each edge;
[0133] Step 5.5, save the vector representations of the tokens, the feature vectors of the nodes of all subgraphs of Sample Three, and the vector features of each edge to Sample Four; Sample Four is a JSON format file.
[0134] Step 6.1, randomly divide the JSON format file of Sample Four into a training set, a validation set, and a test set according to 8:1:1;
[0135] Step 6.2.1, read and parse the data of the training set and the validation set, divide the data into multiple batches according to the set batch size, and convert them into batch graph objects suitable for processing by the graph neural network model;
[0136] Step 6.2.2, when initializing the graph neural network model, it receives the input feature dimension, output feature dimension, number of message passing steps, and maximum number of edge types as parameters. Through these parameters, the model initializes the gated graph convolutional layer and the fully connected layer;
[0137] Step 6.2.3, during the training process, the initialized model performs forward propagation, extracts the graph structure, node features, and edge types from the input batch graph objects, and moves this data to the GPU for accelerated calculation;
[0138] Step 6.2.4, the graph neural network model performs message passing through the gated graph convolutional layer to update the features of each node. Then, the sum of the features of all nodes of each graph is calculated to obtain the global feature representation of the graph. The global feature is mapped to a scalar value (representing the classification score) through the fully connected layer, and finally the sigmoid function is used to compress the output to the range of [0,1] for the final result output of the binary classification task (for example, determining whether there is a vulnerability);
[0139] Step 6.2.5, after completing the forward propagation, calculate the loss value and perform backpropagation to obtain the gradient. Use an optimizer (such as Adam) to update the model parameters. Every certain number of steps (such as 50 steps), the system will enter the validation phase. During the validation phase, calculate the loss, F1 score, and accuracy on the validation set, and record these metrics for subsequent analysis;
[0140] Step 6.2.6, if the F1 score on the validation set is greater than 50 and the accuracy is higher than the historical best value, save the current model weights as the best model. When the performance on the validation set does not improve, the patience counter is incremented. Once the patience counter reaches the pre-set maximum value, the early stopping mechanism is triggered to terminate the training early to avoid overfitting.
[0141] Step 6.3, during the evaluation process, load the best weights pre-trained for the model, and then perform the performance evaluation of vulnerability detection on the test set; the accuracy of the experimental vulnerability detection is 89.24%;
[0142] Control example
[0143] Directly jump the DOT file of Sample 1 in Step 3.1 to Step 5.1, traverse all nodes in Sample 1, and split the code string of each node into lexical tokens (Tokens), such as operators, variable names, keywords, etc. The remaining steps are the same as those in the embodiment. Finally, check the accuracy of vulnerability detection, and the accuracy is 53.32%;
[0144] It can be seen from this that the accuracy of the embodiment of the present invention is 67.37% higher than that of the control example. This significant improvement makes the present invention more suitable for the rapid vulnerability detection of large-scale code libraries, demonstrating superior engineering application value and practical operability.
[0145] The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Those of ordinary skill in the relevant technical fields can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention, and the patent protection scope of the present invention is defined by the claims.
[0146] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A source code program dependency graph pruning method for vulnerability detection, characterized in that: The steps include: Step 1, downloading a source code data set, recording the source code data set as data set A, and performing attribute screening on the data set A to obtain a filtered data set, recording the filtered data set as data set B; Step 2, preprocessing the data set B to remove irrelevant information in the data set B to form sample 1; Step 3, converting the sample 1 into a program dependency graph representation to form a sample 2; the sample 2 can comprehensively present the semantic logic, grammatical features and program structure characteristics of the code of the sample 2 through the joint data structure of control dependency, data dependency and abstract syntax tree; Step 4: Prune the pointer nodes, array nodes, sensitive API call nodes, and integer overflow expression nodes in sample 2 to generate a program dependency subgraph recorded as sample 3. Step 5: Based on the word vector Word2Vec model, the code semantic unit in sample 3 is vectorized and trained to generate a code fragment vector representation that integrates grammatical and semantic features, which is recorded as sample 4. Step 6: Use the graph neural network to train and predict sample 4 to obtain the vulnerability detection results of the source code.
2. The method according to claim 1, characterized in that The step 1 is specifically implemented according to the following steps: Step 1.1, download the source code dataset file containing multiple attributes, denoted as dataset A; Step 1.2, filter out data with the attribute vul field value of 1 from data set A, and construct vulnerability data set B; concatenate the project name project and the commit version number commit_id in the data set B to generate a unique file identifier in the format of {project}_{commit_id}; Step 1.3, storing the vulnerability code of each sample in data set B as a vulnerability file, and storing the repair code of each sample in data set B as a repair file; Step 1.4, by comparing the vulnerable file and the repair file line by line, extracting the difference line numbers between the compared vulnerable file and the repair file, and generating a marking file with the difference line numbers.
3. The method according to claim 2, characterized in that The step 1.4 is specifically implemented according to the following steps: Step 1.4.1: Use a comparison tool to compare the vulnerable file and the repaired file line by line to generate a difference list. If a line starts with a "-" sign, it means that it only exists in the deleted line of the vulnerable file. If a line starts with a "+" sign, it means that it only exists in the added line of the repaired file. Step 1.4.2, in the difference list, if a line is a deleted line beginning with a "-" symbol, the line number of the line is recorded in the annotation file.
4. The method according to claim 1 or 2, characterized in that: The step 2 is specifically implemented according to the following steps: Step 2.1, preprocessing the data set B is: matching the comments in the target file and the repair file by regular expressions to obtain processed source code, wherein the comments include: single-line comment characters and multi-line comment blocks; Step 2.2, create a keyword set, including C language keywords, preprocessor directives, built-in functions and API names; Step 2.3, matching the processed source code through a lexical analyzer and regular expression rules to obtain function names and variable names; Step 2.4, for function names that are not in the keyword set, generate new symbols in the order they appear in the code and replace the original function names; Step 2.5, if the variable name is not in the keyword set and is not the number of command line parameters or the command line parameter vector, then generate new symbols for the variable names that are not in the keyword set in the declaration order to obtain sample one.
5. The method according to claim 1, characterized in that The step 3 is specifically implemented according to the following steps: Step 3.1, parsing sample 1 into a binary file, and converting the binary file into files in DOT and JSON formats; Step 3.2, extract all edges from the DOT file, extract all nodes from the JSON file, filter out nodes that do not contain row number attributes, and form filtered nodes; Step 3.3, sort the filtered nodes according to the column number and line number of the code in the program to form sorted nodes; match and integrate the sorted nodes with all edges through the abstract syntax tree, data dependency graph and control dependency graph to construct a program dependency graph, which is recorded as sample 2.
6. The method according to claim 1 or 2, characterized in that: The step 4 is specifically implemented according to the following steps: Step 4.1, obtain the pointer node, array node, sensitive API call node and integer overflow expression node from sample 2 to obtain the focus of the code slice; Step 4.2, after obtaining the focus, by traversing the parent node in the abstract syntax tree, combining the front and back slicing methods, a program slice list consistent with the actual order of the code is generated; Step 4.3, traverse each subgraph in the program slice list to perform vulnerability check and obtain the subgraph after the check; Perform vulnerability check on the annotated files in step 1.4 above; Step 4.4, for the subgraph after checking in the above step 4.3, initialize a directed graph object, and create a corresponding node in the graph for each node in the subgraph after checking; then, traverse the edge information of each node, record the type of edge and the connected nodes, and add them to the directed graph to obtain the processed directed graph; Step 4.5, verify the edges in the processed directed graph to ensure that the source node and the target node of each edge exist, and generate a final subgraph, which is recorded as sample three.
7. The method according to claim 6, characterized in that The step 4.2 is specifically implemented according to the following steps: Step 4.2.1, for each code concern, perform the following structured traversal operations: first, obtain the direct parent node of the concern in the abstract syntax tree; if the current parent node has a data dependency edge or a method definition node, add it directly to the concern list; otherwise, traverse upward along the abstract syntax tree to the first ancestor node that meets the above conditions; perform deduplication detection on the ancestor node, and if it does not exist in the concern list, dynamically append it to form a new concern list; Step 4.2.2, process all nodes in the new focus list in a loop; for each node in the new focus list, slice it in a backward and forward manner to obtain the successor and predecessor nodes of the node, and add the successor and predecessor nodes to the slice list for integration; Step 4.2.3, when integrating the slice list, exclude unimportant or incomplete successor and predecessor nodes, and sort all successor and predecessor nodes according to line numbers to generate a program slice list that is consistent with the actual order of the code.
8. The method according to claim 1, characterized in that The step 5 is specifically implemented according to the following steps: Step 5.1, traverse the node information of the sample three subgraphs, and split the code string of each node of the sample three subgraph into vocabulary tokens; Step 5.2, training the vocabulary tokens according to the Word2Vec model to generate vector representations of the tokens; Step 5.3, convert sample three into a graph structure, and generate feature vectors of nodes of all sample three subgraphs; Step 5.4, traverse the edge information of sample three and obtain the vector features of each edge; Step 5.5, save the vector representation of the mark, the feature vectors of the nodes of all sample three subgraphs, and the vector features of each edge in sample four.