Vulnerability detection method based on compressed code attribute graph and edge type differentiation processing

By compressing the code attribute graph and differentiating edge types, the problems of low code attribute graph structure efficiency and indistinguishable edge type processing are solved, and efficient and accurate vulnerability detection is achieved.

CN120805137APending Publication Date: 2025-10-17DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510747066.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing code attribute graph processing methods have problems such as low structural efficiency and indiscriminate edge type processing, which leads to limited vulnerability detection accuracy.

Method used

A vulnerability detection method based on compressed code attribute graph and edge type differentiation processing is adopted. Node redundancy is reduced through grammatical subtree aggregation, and independent edge type training functions and attention mechanisms are designed to optimize information flow.

Benefits of technology

The accuracy and efficiency of vulnerability detection have been improved, with the number of nodes reduced by 41.5%, the detection accuracy reaching 99.97%, and the F1 score reaching 99.98%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805137A_ABST
    Figure CN120805137A_ABST
Patent Text Reader

Abstract

The invention discloses a vulnerability detection method based on a compressed code attribute graph and edge type differentiation processing. The method comprises the following steps: automatically extracting the code attribute graph; aggregating the compressed code attribute graph based on the syntax sub-tree; and carrying out edge type differentiation processing on the learning network. According to the method, corresponding nodes of each code attribute graph in an abstract syntax tree are positioned, and a hierarchical gating aggregation method is adopted to compress sub-trees of the syntax tree. The grammar sub-tree aggregation process significantly reduces nodes in a code attribute graph, and the number of graph nodes is reduced to 41.5% of the original number in an open-source software guarantee reference data set. According to the method, an independent training function is designed for each edge type in the compressed code attribute graph, so that differentiation processing of the edge types is realized; meanwhile, an edge type attention mechanism is introduced, the detection accuracy is guaranteed while the graph scale is reduced, the test accuracy of a reference data set is guaranteed to reach 99.97% in open-source software, and the F1 score reaches 99.98%.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to information security detection technology, in particular to a vulnerability detection method based on compressed code attribute graph and edge type differential processing. BACKGROUND

[0002] With the rapid development of information technology, software systems have been widely used in financial, medical, national defense and other key fields, and the security of software is directly related to national security and public interest. However, in the software development process, code vulnerabilities caused by human negligence or design defects may be exploited by malicious attackers, resulting in data leakage, service interruption and other serious consequences. It is crucial to effectively prevent these risks and improve the precision and efficiency of vulnerability detection technology.

[0003] The current mainstream vulnerability detection technology usually processes the code as a graph structure for analysis. In the graph structure of the code, the code attribute graph contains the most abundant semantics. However, the traditional code attribute graph can describe the program structure and data flow, but still has limitations: first, the existing code attribute graph processing method has structural efficiency defects. In the traditional code attribute graph, the leaf nodes of the abstract syntax tree (such as variable names, constant values) occupy a large proportion, but these leaf nodes are usually irrelevant to the subsequent control flow graph or data dependency graph analysis. This significant node redundancy increases the computational burden and reduces the model efficiency. Second, the existing vulnerability detection model based on graph attention network treats all edge types (data dependency, control dependency, control flow) indiscriminately, making it difficult to capture key vulnerability patterns. Different edge types will interfere with each other and become noise, limiting the precision of vulnerability detection. SUMMARY

[0004] To solve the problem of low efficiency of code attribute graph structure, the present application proposes a vulnerability detection method based on compressed code attribute graph and edge type differential processing, which can reduce node redundancy while preserving sub-tree structure information and improve the precision of vulnerability detection.

[0005] The vulnerability detection method based on compressed code attribute graph and edge type differential processing includes the following steps:

[0006] A. Automatic extraction of code attribute graph

[0007] A1. Normalize the code: normalize the code, remove comments, and replace variable names and function names.

[0008] A2. Generate graph structure: batch extract the normalized code attribute graph, which includes the abstract syntax tree, program dependency graph and control flow graph converted by the attribute graph. The three graph structures share a set of nodes.

[0009] A3. Initialize node features: Encode each node using a pre-trained model.

[0010] B. Aggregation and compression of code attribute graph based on grammar subtree

[0011] B1. Locating the root node

[0012] For each node n in the control flow graph and program dependency graph r , find the corresponding node n in the abstract syntax tree u , n u As the root node of the subtree to be compressed, where r is the unique index of the node in the control flow graph and the program dependency graph, and u is the unique index of the node in the abstract syntax tree corresponding to the node in the control flow graph and the program dependency graph.

[0013] B2. Enhanced node features

[0014] For each node n i Inject two codes, the formula is as follows:

[0015]

[0016] Among them, i is the unique index of the child node in the subtree during each round of recursive operation, f i For child node n i The original characteristics, For node n i Enhanced features, pos i For node n i Position index in the sibling list, depth i Represents node n i The depth in the subtree, APE is the absolute position code, recording the order of sibling nodes, RDE is the relative depth code, recording the parent-child relationship, the calculation formula is as follows:

[0017]

[0018]

[0019] Where k is the shard index of the feature dimension, 0≤k≤d / 2, d represents the feature dimension, MLP represents the multi-layer perceptron, and θ is the training parameter set of the multi-layer perceptron.

[0020] B3. Calculate child node weights

[0021] According to the parent node features, enhanced child node features and position index, the aggregation weight g of the child node is calculated through the gating function i , the calculation formula is as follows:

[0022]

[0023] where, is the enhanced node feature, f i is the original feature of the child node, pos i is the node n i , j is the unique index of the parent node in the sub-tree at each round of recursive operation, f j is the feature of the parent node n j , Sigmoid is the Sigmoid activation function, || represents the concatenation operation, W gate is the gating weight matrix, b gate is the bias term, gate represents that the gating weight matrix and the bias term are only the parameters of the current gating function, and φ is the sequential mapping function used to normalize the position index, which is defined as follows:

[0024]

[0025] where, n j represents the parent node, is the set of child nodes of n j .

[0026] B4, aggregating child node features

[0027] The parent node n j feature is obtained by bottom-up hierarchical weighted aggregation. The aggregation method is as follows:

[0028]

[0029] where, represents the historical aggregated feature, is the enhanced node feature, f j is the original feature of the parent node, is the set of child nodes of n j , g i is the gating value of the node n i , ReLU is the rectified linear unit, W a is the aggregation weight matrix, b a is the bias term, a represents that the aggregation weight matrix and the bias term are only the parameters of the current aggregation function.

[0030] The root node feature of the sub-tree is finally obtained , that is, the feature of each node in the compressed code attribute graph is obtained.

[0031] B5, compressing the code attribute graph

[0032] After aggregation, the feature vector of the root node already contains the information of the sub-tree, and finally the original sub-tree only retains the aggregated root node.

[0033] C. Edge type differentiation processing learning network

[0034] After the code attribute graph is compressed, the vulnerability detection accuracy is further improved through edge type differentiation.

[0035] C1. Differentiated edge type processing

[0036] Take a pair of connected nodes n in the graph obtained after code attribute graph compression u and n v , v represents the node n u The unique index of the connected node, node n u and n v The edge between t represents the tth edge type, then the feature transformation of each edge type is calculated as follows:

[0037]

[0038] in, represents the edge node interaction features after feature transformation, and Is the edge type e t The training weights and biases of and Represents node n u and n v The feature of is represented by , and ReLU is a rectified linear unit. Through feature transformation, each edge type is adapted to an independent learning mechanism to optimize information transfer between nodes.

[0039] C2. Edge-type attention mechanism

[0040] For each edge e t Calculating attention weights Dynamically adjust node n u With neighbors v The information flow weight between them is calculated as follows:

[0041]

[0042] in, Represents node n u The set of neighbor nodes, n v Indicates n u A neighbor node, n w It is a temporary variable used to traverse all neighbors in the sum operation, and w is the unique index of the temporary variable. and Represents node n u 、n v and n w Features, Is the edge type et learning parameter matrix, LeakyReLU is a leaky rectified linear unit, and formula (9) calculates the edge type e t For node n u and n v , the information transmission effect is dynamically adjusted according to the type of the edge.

[0043] C3, updating the node feature

[0044] After the edge type differentiation processing and the attention mechanism, finally, the feature vector of each node n u is updated according to the features of the neighbor nodes n v and the weight of the edge, and the updated node feature is calculated by the following formula:

[0045]

[0046] wherein, represents the edge node interaction feature obtained by feature transformation, represents the attention weight of the edge e t between n u and the neighbor n v , and represent the features of the nodes n u and n v , represents the neighbor node set of the node n u , represents the element-wise product, W node and b node are training weights and biases, and represents the training weights and biases in formula (10) are only parameters in the node updating process, ReLU is a linear rectifier unit, which ensures that each node dynamically integrates information from different types of edges and adjusts the information flow through the edge type attention mechanism.

[0047] C4, vulnerability classification

[0048] The global feature representation of the graph is generated by averaging the feature vectors of all nodes using global average pooling, and the graph-level feature vector is obtained The calculation formula is as follows:

[0049]

[0050] wherein, N is the number of nodes in the graph.

[0051] The graph-level feature is passed to the classifier to obtain the classification result, and the calculation formula is as follows:

[0052]

[0053] wherein, W class and b class are the training weights and bias of the classifier, class represents the training weights and bias in the formula are only parameters of the classifier, Sigmoid is a sigmoid activation function, is a predicted probability value, If the code contains a vulnerability.

[0054] The loss function adopts binary cross-entropy, and the calculation formula of the loss function is as follows:

[0055]

[0056] wherein, y is an actual label, is a predicted label, minimizing the loss function optimizes the classifier parameters and improves the detection accuracy.

[0057] Compared with the prior art, the beneficial effects of the present application are as follows:

[0058] 1. Ensure the accuracy of the classification task while improving the structural efficiency of the code property graph processing.

[0059] In the prior art code property graph processing method, the leaf nodes (such as variable names and constant values) of the abstract syntax tree occupy most of the graph, however, most of these leaf nodes are irrelevant to the subsequent control flow graph and data dependency graph analysis. Therefore, the traditional method will cause the graph neural network to process a large number of invalid nodes, increase the computational burden, and reduce the processing efficiency. The present application proposes a code property graph compression method based on syntax subtree aggregation, locates the corresponding node of each code property graph in the abstract syntax tree, and adopts a hierarchical gating aggregation method to compress the syntax tree subtree. The syntax subtree aggregation process significantly reduces the nodes in the code property graph, and reduces the number of graph nodes in the open source software security reference dataset to 41.5% of the original.

[0060] 2. Improve the adaptability of the graph neural network in code analysis

[0061] The existing method for vulnerability detection based on graph attention network usually uniformly processes all types of edges (such as data dependency edges, control dependency edges, and control flow edges), and fails to consider the characteristics of different edge types in vulnerability detection. The present application designs independent training functions for each edge type in the compressed code property graph, realizes the differentiated processing of edge types, and at the same time, introduces an edge type attention mechanism to dynamically adjust the weight of the information flow according to the characteristics of each edge type, which reduces the size of the graph while ensuring the accuracy of the detection, and the test accuracy on the open source software security reference dataset reaches 99.97%, and the F1 score reaches 99.98%. BRIEF DESCRIPTION OF DRAWINGS

[0062] The present invention has the following attached drawings Figure 5 Zhang, wherein:

[0063] Figure 1 is a code example diagram.

[0064] Figure 2 is a code attribute diagram.

[0065] Figure 3 is a framework diagram of the present invention.

[0066] Figure 4 is an experimental result of the present invention.

[0067] Figure 5 is a comparative experimental result. DETAILED DESCRIPTION

[0068] The present invention will be further described below in conjunction with the drawings.

[0069] Figure 1 is a code example diagram, Figure 1 the code attribute diagram of the code shown in Figure 2 is shown, and the code attribute diagram is obtained by attribute graph conversion and combination of the abstract syntax tree, the control flow graph and the program dependency graph.

[0070] Figure 3 is a framework diagram of a vulnerability detection method based on compressed code attribute graph and edge type differential processing. First, the source code is normalized to generate the corresponding code attribute graph. Then, the nodes corresponding to the abstract syntax tree are located through the program dependency graph and the control flow graph, and the located nodes are taken as the aggregation root nodes of the syntax sub-tree. In the sub-tree aggregation process, the nodes are aggregated layer by layer from the leaf node to the root node, while recording the relative depth information of each node and its position order in the sibling nodes, and finally a compressed code attribute graph containing the semantic information and structural information of the abstract syntax tree is generated. The syntax sub-tree aggregation process significantly reduces the nodes in the code attribute graph, and in the open source software security reference data set, the number of graph nodes is reduced to 41.5% of the original, and the node change statistics are shown in Table 1.

[0071] Table 1: Comparison of node numbers before and after compression of code attribute graph

[0072]

[0073] The compressed graph contains three types of edges, namely control flow edges, data dependency edges and control dependency edges. Independent edge feature processing mechanisms are designed for different types of edges. By introducing an attention mechanism, the weight of each type of edge is dynamically calculated. Based on the calculated edge weight, the features of adjacent nodes are weighted and aggregated to update the node representation. Next, the feature vectors of all nodes are input into the global average pooling layer to extract the graph-level semantic features, and the obtained graph representation is input into the classifier for training and reasoning, and finally the vulnerability detection result is output.

[0074] Figure 4 For the experimental results of the vulnerability detection method based on compressed code attribute graph and edge type differential processing, the loss, accuracy and F1 score of the vulnerability model based on compressed code attribute graph and edge type differential processing in the training process are shown. The trend of the number of rounds, and each index is compared with the performance of the training set and the validation set. Figure 4 The left subgraph in the middle is the training and validation loss curve. The initial training loss and validation loss are high, about 0.5. The loss decreases rapidly in the first 10 rounds of training, indicating that the model learns quickly. Then the loss value tends to be stable, close to 0, indicating that the model has reached a good fit. The validation loss and the training loss have similar trends, and there is no obvious upward trend, indicating that there is no overfitting. Figure 4 The middle subgraph is the training and validation accuracy curve. After the 5th round, the training and validation accuracy quickly rises above 0.95; then both tend to be close to 1, indicating that the model performs well on both sets of data; the validation accuracy is almost the same as the training accuracy, indicating that the model has strong generalization ability. Figure 4 The right subgraph is the training and validation F1 score curve. The F1 score is low at the beginning, but rises quickly and tends to 1. The F1 score of the validation set and the training set is almost synchronous, indicating that the model not only has high precision, but also has good recall. The F1 score has almost no fluctuation in the later training period, indicating that the model is stable and reliable.

[0075] Figure 5 For the experimental results of the vulnerability detection method based on compressed code attribute graph and not performing edge type differential processing, Figure 5 The experimental environment is consistent with that of Figure 4 The experimental environment is consistent with that of Figure 5 The left subgraph is the training and validation loss curve. There is a clear fluctuation in the middle of the curve, and the validation loss fluctuates greatly after the 15th round. The model is not stable during the training process. Figure 5 The middle subgraph is the training and validation accuracy curve. The accuracy rises rapidly in the first ten rounds, and then remains at the level of 0.95, but there is a clear fluctuation, indicating that the model will have inconsistent performance on the validation set. Figure 5The right subgraph is a training and verification F1 score curve, and the F1 score fluctuates at a high level, indicating that the method without edge type differentiation processing has fluctuation in the recognition ability of positive and negative samples in different rounds.

[0076] By Figure 4 And Figure 5 It can be seen that, under the same data and training batch, the network training process using edge type differentiation processing is more stable, and the performance is significantly improved, Figure 4 All indicators in the middle period decrease first, then rapidly increase and stabilize, and the difference between training and verification is small, indicating that edge type differentiation processing has better stability and generalization ability.

[0077] The present application is not limited to the present embodiment, any equivalent concept or change within the technical scope disclosed in the present application is included in the protection scope of the present application.

Claims

1. A vulnerability detection method based on compressed code attribute graph and edge type differentiation processing, characterized by: The following steps are involved: A. Automatically extract code attribute graph; B. Aggregate and compress code attribute graph based on grammatical subtree; C. Edge type differentiation processing learning network.

2. The vulnerability detection method based on compressed code attribute graph and edge type differentiation processing according to claim 1 is characterized by: The method for automatically extracting the code attribute graph described in step A includes the following steps: A1. Standardize the code: standardize the code, remove comments, and replace variable and function names; A2. Generate graph structures: Batch extract normalized code attribute graphs. The extracted code attribute graphs contain the abstract syntax tree, program dependency graph, and control flow graph that have been converted through the attribute graph. These three graph structures share a set of nodes. A3. Initialize node features: Encode each node using a pre-trained model.

3. The vulnerability detection method based on compressed code attribute graph and edge type differentiation processing according to claim 1 is characterized by: The method of aggregating and compressing a code attribute graph based on a grammar subtree in step B comprises the following steps: B1. Locating the root node For each node n in the control flow graph and program dependency graph r , find the corresponding node n in the abstract syntax tree u , n u As the root node of the subtree to be compressed, where r is the unique index of the node in the control flow graph and the program dependency graph, and u is the unique index of the node in the abstract syntax tree that corresponds to the node in the control flow graph and the program dependency graph; B2. Enhanced node features For each node n i Inject two codes, the formula is as follows: Among them, i is the unique index of the child node in the subtree during each round of recursive operation, f i For child node n i The original characteristics, For node n i Enhanced features, pos i For node n i Position index in the sibling list, depth i Represents node n i The depth in the subtree, APE is the absolute position code, recording the order of sibling nodes, RDE is the relative depth code, recording the parent-child relationship, the calculation formula is as follows: Where k is the shard index of the feature dimension, 0≤k≤d / 2, d represents the feature dimension, MLP represents the multi-layer perceptron, and θ is the set of training parameters of the multi-layer perceptron. B3. Calculate child node weights According to the parent node features, enhanced child node features and position index, the aggregation weight g of the child node is calculated through the gating function i , the calculation formula is as follows: in, is the enhanced node feature, f i is the original feature of the child node, pos i For node n i The position index in the sibling list, j is the unique index of the parent node in the subtree during each round of recursive operation, f j is the parent node n j The features of , Sigmoid is the S-type activation function, || represents the splicing operation, W gate is the gating weight matrix, b gate is the bias term, gate indicates that the gating weight matrix and the bias term are only the parameters of the current gating function, and φ is the sequential mapping function used to normalize the position index, which is defined as follows: Among them, n j Represents the parent node, n j The collection of child nodes; B4. Aggregation sub-node features Bottom-up hierarchical weighted aggregation to get parent node n j feature The aggregation method is as follows: in, represents historical aggregation features, is the enhanced node feature, f j is the original feature of the parent node, n j The child node set of g i For node n i The gate value, ReLU is the rectified linear unit, W a is the aggregation weight matrix, b a is the bias term, a indicates that the aggregation weight matrix and the bias term are only the parameters of the current aggregation function; Finally, the root node characteristics of the subtree are obtained That is, the characteristics of each node in the compressed code attribute graph are obtained; B5. Compression code attribute diagram After the aggregation is completed, the feature vector of the root node already contains the information of the subtree, and finally the original subtree only retains the root node after aggregation.

4. The vulnerability detection method based on compressed code attribute graph and edge type differentiation processing according to claim 1 is characterized by: The method for learning a network by differentiating edge types in step C includes the following steps: C1. Differentiated edge type processing Take a pair of connected nodes n in the graph obtained after code attribute graph compression u and n v , v represents the node n u The unique index of the connected node, node n u and n v The edge between t represents the tth edge type, then the feature transformation of each edge type is calculated as follows: in, represents the edge node interaction features after feature transformation, and Is the edge type e t The training weights and biases of and Represents node n u and n v The feature of , ReLU is the rectified linear unit; through feature transformation, each edge type is adapted to an independent learning mechanism, thereby optimizing the information transfer between nodes; C2. Edge-type attention mechanism For each edge e t Calculating attention weights Dynamically adjust node n u With neighbors v The information flow weight between them is calculated as follows: in, Represents node n u The set of neighbor nodes, n v Indicates n u A neighbor node, n w It is a temporary variable used to traverse all neighbors in the sum operation, and w is the unique index of the temporary variable. and Represents node n u 、n v and n w Features, Is the edge type e t The learning parameter matrix of LeakyReLU is the leaky rectified linear unit, and formula (9) calculates the edge type e t For node n u and n v The information transmission effect is calculated and the propagation weight of the information is dynamically adjusted according to the type of edge; C3. Update node features After edge type differentiation and attention mechanism, each node n u The feature vector of the neighbor node n v The features and edge weights are updated, and the updated node features Calculated by the following formula: in, represents the edge node interaction features obtained by feature transformation, Indicates n u and neighbors v The edge e t The attention weight, and Represents node n u and n v Features, Represents node n u The set of neighbor nodes, ⊙ represents the element-wise product, W node and b node are the training weights and biases, indicating that the training weights and biases in formula (10) are only parameters in the node update process. ReLU is a linear rectified unit, which ensures that each node dynamically integrates information from different types of edges and adjusts the information flow through the edge type attention mechanism. C4. Vulnerability Classification Use global average pooling to average the feature vectors of all nodes to generate the global feature representation of the graph and obtain the graph-level feature vector The calculation formula is as follows: Where N is the number of nodes in the graph; Graph-level features Pass it to the classifier to get the classification result. The calculation formula is as follows: Among them, W class and b class are the training weights and biases of the classifier. Class indicates that the training weights and biases in the formula are only the parameters of the classifier. Sigmoid is the S-type activation function. is the predicted probability value, The code is considered to contain vulnerabilities; The loss function uses binary cross entropy, the loss function The calculation formula is as follows: Among them, y is the actual label, To predict labels, minimize the loss function and optimize the classifier parameters to improve detection accuracy.