Method for carrying out code bug repair detection by using graph embedded component to form feature graph structure

The feature graph structure is constructed through the graph embedding component, which solves the problems of manual analysis and static rule dependence in code vulnerability repair detection in the existing technology, and achieves a comprehensive understanding and feature extraction of the vulnerability repair process, which significantly improves the accuracy and efficiency of detection.

CN120068083APending Publication Date: 2025-05-30DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510070416.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology has problems of manual analysis and static rule dependence in code vulnerability repair detection, which leads to high resource consumption and low detection accuracy, making it difficult to capture global code dependencies and structured information, and the false detection and missed detection rates are high.

Method used

The graph embedding component is used to build a feature graph structure. By modeling the complex dependencies and repair processes of the code into graphs, and extracting key features in combination with graph embedding technology, we can achieve a comprehensive understanding of the vulnerability repair process and feature extraction.

Benefits of technology

It significantly improves the accuracy and efficiency of vulnerability repair detection, reduces the false detection and missed detection rates, can more effectively capture the complex relationships and global features of the code, and improves the comprehensiveness and accuracy of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068083A_ABST
    Figure CN120068083A_ABST
Patent Text Reader

Abstract

The invention provides a method for performing code bug repair detection by forming a feature graph structure by utilizing a graph embedding component. The method comprises the following steps: S1, obtaining a target data set after bug repair; s2, representing the text feature codes as grammar and grammar information of the text by using a graph embedding component, and capturing a feature structure graph of text keywords; s3, non-text features of the target data set are obtained, the non-text features are regarded as independent nodes, the text feature graph and non-text feature nodes are aggregated into a text and non-text aggregation graph, and code information and similarity information on each node of the text and non-text aggregation graph are obtained; s4, performing similarity calculation on the code information and the similarity information in the weight enabling component; and S5, inputting node information of the weight-enabled code into a classifier to obtain a detection result. According to the method, the core features of vulnerability repair can be captured at a higher level, so that the accuracy and efficiency of vulnerability repair detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vulnerability repair detection. In particular, it relates to a method for code vulnerability repair detection by using a graph embedding component to form a feature graph structure. Background Art

[0002] Code vulnerability repair is an important part of software development and maintenance. The repair process generally includes vulnerability detection, analysis, solution formulation, implementation of repair, test verification, deployment and going live, etc. Traditional vulnerability repair detection methods often rely on manual analysis or static rule-based methods, and it is difficult to effectively process high-dimensional and complex data features. In addition, with the growth of software code scale, the complexity and diversity of vulnerabilities have also increased significantly, which poses higher requirements on existing detection tools.

[0003] In recent years, graph embedding technology has shown unique advantages in processing complex structured data. For example, models such as GraphSAGE perform node embedding by sampling local subgraphs, which can effectively process large-scale graph data and are suitable for extracting local structural features from code dependency graphs. In addition, GAT (Graph Attention Network) introduces an attention mechanism to dynamically adjust weights according to the importance between nodes and performs well in capturing the dependency relationships of code fragments. GCN (Graph Convolutional Network) aggregates the features of each node with the information of its neighbor nodes by performing convolutional operations on the graph, and can accurately capture the global dependency relationships between different modules of the code. However, although these graph embedding technologies have achieved success in other fields, their applications in the field of vulnerability repair detection are still relatively limited. Current research is more in the exploratory stage. Existing solutions mostly focus on extracting features based on building code dependency graphs or call graphs, but no widely applicable solution has been formed yet. Nevertheless, deep learning technology has been gradually introduced into this field in recent years. For example, by combining graph embedding and natural language processing models (such as CodeBERT) to enhance the understanding of code semantics and structure, it is expected to promote the progress of vulnerability repair detection.

[0004] There are several problems in the existing technology for vulnerability repair detection. First, the methods based on manual analysis or static rules consume a large amount of human and material resources and the detection accuracy is not high enough. Second, the traditional methods have limited ability to capture the global dependencies of the code. Usually, they can only analyze local features and ignore the complex dependencies across modules and files. This leads to the inability to effectively identify the deep associations in the code when detecting vulnerability repairs, thus affecting the comprehensiveness of the detection. There are also deficiencies in the feature expression ability of the existing technology. Although some methods introduce deep learning and attention mechanisms to analyze the syntax and semantics of the code, the ability to extract structured information in code repair, especially the extraction ability of code changes and dependency contexts, is weak. This makes it easy to have false positives or false negatives when dealing with complex code. In addition, although graph embedding technology shows great potential in processing structured data, its application in vulnerability repair detection is still in its infancy. The existing graph structure modeling is not perfect and fails to fully utilize the advantages of graph embedding technology in capturing complex code relationships and extracting global features. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to propose a method for code vulnerability repair detection by using a graph embedding component to form a feature graph structure. By modeling the complex dependencies of the code and the repair process as a graph and combining graph embedding technology to extract key features, this method can capture the core features of vulnerability repair at a higher level, thereby improving the accuracy and efficiency of vulnerability repair detection.

[0006] The technical means adopted by the present invention are as follows:

[0007] A method for code vulnerability repair detection by using a graph embedding component to form a feature graph structure, comprising the following steps:

[0008] S1. Obtain the target data set after vulnerability repair, extract the added and modified code in the target data set after vulnerability repair, and use the added and modified code and the description information of the code as text features;

[0009] S2. Use the graph embedding component to encode and represent the text features as the syntax and grammar information of the text and capture the feature structure graph of the text keywords. The text features are the code information of the vulnerability repair;

[0010] S3. Obtain the non-text features of the target data set, regard the non-text features as separate nodes, and aggregate the text feature graph and the non-text feature nodes into a text and non-text aggregation graph. Each node of the text and non-text aggregation graph represents the code information and the similarity information between the codes;

[0011] S4. In the weight empowerment component, calculate the similarity between the code information and the similarity information to empower the code with weights;

[0012] S5. Input the node information of the weighted code into the classifier to obtain the detection result.

[0013] Further, S2 specifically includes the following steps:

[0014] Deduplicate the words and map each word to an index. Let the sliding window size be K. Starting from the beginning of the sentence, take each group of three words as a sliding window to obtain a set of complete sliding windows.

[0015] Regard each word in each sliding window as a node x i or x j , and calculate the co-occurrence times between nodes x n in the current sliding window W i and x j .

[0016] By adding the co-occurrence times between two nodes, obtain the initial weight x ij of the edge between the two nodes, that is, the text feature map.

[0017] Use the word embedding component to assign an initial value to each word; use the Porter stemming algorithm to obtain the root form of each word, and assign a fixed-length vector representation to each word.

[0018] Apply a word embedding dictionary trained with Glove2, and input the text features into this word embedding dictionary to obtain the embedding vector of each word.

[0019] Further, S3 specifically includes the following steps:

[0020] For non-text features, obtain the pre-modification and post-modification code information from the code before and after vulnerability repair, and use TF-IDF to calculate the scores of tokens in the two code changes.

[0021] Embed all the scores into a vector.

[0022] Obtain two vectors for the pre-modification and post-modification, and use cosine similarity to calculate the similarity value between the two vectors.

[0023] Expand the similarity value of each clue feature so that the similarity value of each clue feature is the same as the embedding length of the word.

[0024] Regard each non-text feature as a node, and connect K - 1 nodes in sequence with the initial edge weight.

[0025] Add a root node and connect the graph based on text features and the graph based on non-text features with the initial edge weight.

[0026] Further, S4 specifically includes the following steps:

[0027] Use a multi-layer perceptron to map the initial values of each node and edge in the text and non-text aggregation graph to initial vectors:

[0028]

[0029] Among them, x i represents the initial node value of a node composed of ten features, and h vi (0) represents the initial node vector after each node v i is mapped, and e ij represents the initial edge vector after the edge feature vector x i and v j between two adjacent nodes v ij is mapped;

[0030] In the propagation layer, the old nodes composed of text features and non-text features are mapped to new node vectors, that is, the node states are updated; the representation of each node vector will accumulate the late information of its local neighborhood through multi-layer propagation:

[0031]

[0032] Among them, f message is a multi-layer perceptron applied to the concatenated input;

[0033] Calculate the attention coefficient through the graph matching network with the nodes in another feature graph:

[0034]

[0035] Among them, f match is a function that transmits cross-feature graph information;

[0036] Use the following formula to calculate the weight from node j to node i:

[0037]

[0038] Among them, s h is the vector space similarity;

[0039] Substitute the attention module into the formula for calculating the attention coefficient, and calculate the cumulative matching value between two prs as follows:

[0040]

[0041] The old vector of a node, the attention coefficients of adjacent nodes in the same feature map, and the attention coefficients of nodes in another feature map are used to update the state of the node. The node update formula is defined as follows:

[0042]

[0043] Among them, f node represents a multi-layer perceptron;

[0044] Update the state of node i according to the attention coefficient from node j to node i and the cumulative matching value between node j and node i.

[0045] After obtaining the graph information after the weight empowerment processing result, input it into the classifier for classification to obtain the detection result of vulnerability repair.

[0046] The present invention also provides a storage medium, which includes a stored program. When the program runs, it executes any one of the above methods for detecting code vulnerability repair by using a graph embedding component to construct a feature map structure.

[0047] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and operable on the processor. The processor runs the computer program to execute any one of the above methods for detecting code vulnerability repair by using a graph embedding component to construct a feature map structure.

[0048] Compared with the prior art, the present invention has the following advantages:

[0049] The core innovation of the present invention lies in the comprehensive integration and deep feature extraction of code changes and vulnerability repair descriptions. Traditional methods, such as VulFixMiner, only focus on code changes and ignore the repair description information, resulting in an incomplete semantic understanding of vulnerability repair. The present invention constructs a feature map structure through graph embedding technology, which can not only accurately extract code changes but also capture important context information from the repair description, realizing a comprehensive understanding of the vulnerability repair process.

[0050] In addition, the present invention also effectively solves the problem of difficult handling of complex dependency relationships in the prior art by means of graph construction, making the vulnerability repair detection perform better in complex projects. At the same time, by combining deep learning and graph embedding technology, the present invention has significant improvements in the accuracy, robustness, and scalability of the model, can effectively reduce the false detection rate and missed detection rate, and improve the overall performance of vulnerability repair detection.

[0051] In summary, by innovatively introducing a graph embedding component, the present invention significantly improves the accuracy and efficiency of vulnerability repair detection, makes up for the deficiencies of the prior art in information integration, complex dependency processing, and feature extraction, and has a broader application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0053] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0056] As Figure 1 shown, the present invention provides a method for detecting code vulnerability repair by using a graph embedding component to form a feature graph structure, including the following steps:

[0057] S1. Obtain a target data set after vulnerability repair, extract the added and modified code in the target data set after vulnerability repair, and use the added and modified code and the description information of the code as text features;

[0058] To evaluate the effectiveness of the model, a new dataset was constructed. The sources of the dataset mainly include:

[0059] The SAP dataset. The SAP dataset contains 1,055 vulnerability fix commits related to 183 Java open-source projects. To ensure the practical usability of these projects, based on the data analysis of SAP, the vulnerability assessment tool Vu-las was run for verification. Subsequently, the data of these vulnerability fix commits were manually collected by monitoring the disclosure of vulnerabilities (not limited to NVD, but also including the web pages of specific projects).

[0060] The second source is all CVEs related to Java and Python disclosed as of January 26, 2021. From the CVEs, 199 commits, 227 issues, and 155 pull requests in Java, and 288 commits, 244 issues, and 353 pull requests in Python were collected. Then, the commits mentioned in the pull requests and issues were extracted. Finally, after removing duplicate commits, all the commits were merged into a single dataset.

[0061] S2. Use the graph embedding component to encode and represent the text features as the syntactic and grammatical information of the text and capture the feature structure diagram of the text keywords. The text features are the description information of the vulnerability fix.

[0062] In the data preprocessing stage, that is, the graph embedding component embedding stage. For text features, in order to utilize the co-occurrence information of global words, a sliding window of a fixed size is used to process the text features on all texts. First, the words are de-duplicated, and each word is mapped to an index. Suppose the sliding window size is set to K. Starting from the beginning of the sentence, each group of three words is used as a sliding window, so as to obtain a set of complete sliding windows. Subsequently, each sliding window is processed as follows. Each word in each sliding window is regarded as a node xi, and the co-occurrence times between the nodes x n in the current sliding window W i and x j are calculated. Then, by adding the co-occurrence times between the two nodes, the initial weight of the edge between the two nodes x ij is obtained. Then, the word embedding component is used to assign an initial value to each word. The Porter stemming algorithm is used to obtain the root form of each word (for example, "works" becomes "work"), and a fixed-length vector representation is assigned to each word, where the length is set to 300. Then, a word embedding dictionary trained using Glove2 is applied to obtain the embedding vector of each word.

[0063] S3. Obtain the non-text features of the target dataset, regard the non-text features as separate nodes, aggregate the text feature map and the non-text feature nodes into a text and non-text aggregation graph, and the code information and similarity information on each node of the text and non-text aggregation graph;

[0064] For non-text features, first obtain the code information before and after modification from two different versions of the code. First, use Term-Frequency Inverse Document Frequency (TF-IDF) to calculate the scores of tokens in these two code changes. Then, embed all the scores into a vector. Finally, obtain two vectors and use cosine similarity to calculate the similarity value between the two vectors.

[0065] After obtaining the similarity value of each clue feature, expand it to 300 dimensions by adding 299 zeros, which is the same as the embedding length of words. Regard each non-text feature as a node, and connect K - 1 nodes in sequence with the initial edge weight. Finally, add a root node and connect the graph based on text features and the graph based on non-text features with the initial edge weight.

[0066] S4. In the weight empowerment component, calculate the similarity between the code information and the similarity information to empower weights to the code;

[0067] In the classification stage of the data, traverse the entire graph structure, calculate the attention coefficients between nodes, and dynamically assign greater weights to the code change nodes with the minimum repeatability through the attention coefficients. First, use a multi-layer perceptron to map the initial values of each node and edge in the feature map to initial vectors:

[0068]

[0069] Among them, x i represents the initial node value of a node composed of ten features, h vi (0) represents the initial node vector after mapping for each node v i , e ij represents the edge feature vector between two adjacent nodes (v i and v j ), and x ij is the initial edge vector after mapping.

[0070] In the propagation layer, the old nodes composed of text features and non-text features are mapped to new node vectors, that is, the node states are updated. Therefore, the representation of each node vector will accumulate the late information of its local neighborhood through multi-layer propagation.

[0071]

[0072] where f message is a multilayer perceptron applied to the concatenated input.

[0073] Secondly, the graph matching network calculates the attention coefficients with the nodes in another feature graph, considering the joint reasoning between the two PRs and solving the problem of lack of joint reasoning between code changes in different versions.

[0074]

[0075] where f match is a function that transfers cross-feature map information. Here, an attention-based module is used.

[0076]

[0077] where s h is the vector space similarity.

[0078] Then, substitute the attention module into the attention coefficient calculation formula to calculate the cumulative matching value between the two pr as follows:

[0079]

[0080] In this way, for an old vector of a node, its attention coefficient with the adjacent nodes in the same feature map and its attention coefficient with the node in another feature map can be used to update the state of the node. The node update formula is defined as follows:

[0081]

[0082] where f node stands for Multilayer Perceptron.

[0083] S5. Input the node information of the code after weight empowerment into the classifier to obtain the detection result.

[0084] In this way, the state of the node can be updated through the attention coefficient, that is, the weight of the node can be empowered, and the graph information after the weight empowerment processing result is obtained and input into the classifier for classification to obtain the detection result of the vulnerability repair.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for code vulnerability repair detection using a graph embedding component to form a feature graph structure, characterized in that: The steps include: S1. Obtain the target dataset after the vulnerability is fixed, extract the codes added and modified in the target dataset after the vulnerability is fixed, and use the added and modified codes and the description information of the codes as text features; S2, using the graph embedding component to encode text features into the grammatical and syntax information of the text and capture the feature structure graph of the text keywords, the text features are the code information of the vulnerability repair; S3, obtaining non-text features of the target data set, treating the non-text features as separate nodes, aggregating the feature structure graph of text keywords and non-text feature nodes into a text and non-text aggregation graph, each node of the text and non-text aggregation graph represents code information and similarity information between codes; S4. In the weight empowerment component, the code information and the similarity information are similarity calculated to weight empower the code; S5. Input the node information of the code after weight empowerment into the classifier to obtain the detection result.

2. The method for code vulnerability repair detection using a graph embedding component to form a feature graph structure according to claim 1, characterized in that: S2 specifically includes the following steps: Remove duplicate words and map each word to an index. Set the sliding window size to K. Starting from the beginning of the sentence, use each group of three words as a sliding window to obtain a complete set of sliding windows. Consider each word in each sliding window as a node x i or x j , calculate the current sliding window W n Midpoint x i and x j The number of co-occurrences between The initial weight x of the edge between two nodes is obtained by adding the number of co-occurrences between the two nodes before ij , i.e., text feature map; Use the word embedding component to assign an initial value to each word; use Porter stemming to obtain the root form of each word and assign a fixed-length vector representation to each word; Apply a word embedding dictionary trained using Glove2 and input the text features into this word embedding dictionary to obtain the embedding vector for each word.

3. The method for code vulnerability repair detection using a graph embedding component to form a feature graph structure according to claim 1, characterized in that: S3 specifically includes the following steps: For non-text features, we obtain the code information before and after the vulnerability is fixed, and use TF-IDF to calculate the scores of the tokens in the two code changes. Embed all scores into a vector; Get the two vectors before and after the modification, and use cosine similarity to calculate the similarity value of the two vectors before and after the modification; Expand the similarity value of each clue feature so that the similarity value of each clue feature is the same as the embedding length of the word; Treat each non-text feature as a node and connect K-1 nodes in sequence using the initial edge weights; Add a root node and connect the graph based on text features and the graph based on non-text features with initial edge weights.

4. The method for code vulnerability repair detection using a graph embedding component to form a feature graph structure according to claim 1, characterized in that: S4 specifically includes the following steps: Use a multilayer perceptron to map the initial values ​​of each node and edge in the text and non-text aggregate graph to an initial vector: Among them, x i represents the initial node value of a node consisting of ten features, h vi (0) Represents each node v i The initial node vector after mapping, e ij Represents two adjacent nodes v i and v j The marginal feature vector x between ij The initial edge vector after mapping; In the propagation layer, the old nodes composed of text features and non-text features are mapped to new node vectors, that is, the node status is updated; the representation of each node vector will accumulate the later information of its local neighborhood through multi-layer propagation: Among them, f message is a multilayer perceptron applied to the concatenated input; The attention coefficient is calculated by matching the graph network with nodes in another feature graph: Among them, f match It is a function that transfers cross-feature map information; Use the following formula to calculate the weight from node j to node i: Among them, s h is the vector space similarity; Substituting the attention module into the attention coefficient calculation formula, the cumulative matching value between the two pr is calculated as follows: The old vector of a node is used to update the state of the node with the attention coefficient of the adjacent nodes in the same feature map and the attention coefficient of the node in another feature map. The node update formula is defined as follows: Among them, f node stands for Multilayer Perceptron; Update the state of node i according to the attention coefficient from node j to node i and the cumulative matching value between node j and node i; After obtaining the graph information after the weight empowerment processing result, it is input into the classifier for classification to obtain the detection result of the vulnerability repair.

5. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the method for detecting code vulnerability repair by using a feature graph structure constructed by using a graph embedding component as described in any one of claims 1 to 4 is executed.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the method for code vulnerability repair detection by using a graph embedding component to construct a feature graph structure as described in any one of claims 1 to 4 through the computer program.