Source code vulnerability slice-level detection and statement-level positioning method based on multi-task learning

Through the method of multi-task learning and graph neural network combined with Transformer model, the problem of difficulty in accurately positioning and capturing long dependencies in the prior art is solved, and vulnerability detection with higher accuracy and fine-grained size is achieved.

CN120180438APending Publication Date: 2025-06-20NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311757238.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the vulnerability detection, it is difficult to detect code vulnerabilities at a higher granularity, and it is impossible to give specific statements that may lead to vulnerabilities. The graph-based model is limited by calculation costs when capturing long-range dependencies, and the detection accuracy is insufficient.

Method used

The source code vulnerability detection method based on multi-task learning is adopted to capture vulnerability features through program slicing technology, and combined with graph neural network and Transformer model, multi-task learning for whole graph classification and node classification is modeled to improve the model's ability to capture long dependencies.

Benefits of technology

It realizes more accurate detection of vulnerabilities in the source code at the slice level and statement level, improves detection accuracy and fine-graining, and can better capture vulnerability features and long dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180438A_ABST
    Figure CN120180438A_ABST
Patent Text Reader

Abstract

The invention relates to a source code vulnerability slice-level detection and statement-level positioning method based on multi-task learning, and belongs to the technical field of source code detection. The invention provides a new source code vulnerability detection method based on deep learning. According to the method, the program slice sub-graph related to the vulnerability candidate elements is extracted from the whole source code based on the accessibility of the program dependency graph, and vulnerability detection is carried out by taking the slice sub-graph as granularity, so that vulnerability features can be better captured, and the detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting source code vulnerability slicing level and statement level positioning based on multi-task learning, belonging to the technical field of source code detection. Background Art

[0003] Static vulnerability detection is a key technology in the software development process, aiming to discover and eliminate vulnerabilities in the early stage of development. Traditional static vulnerability detection tools rely on security experts to analyze vulnerability patterns and formulate detection rules. Such a process is time-consuming, laborious and error-prone. With the growth of software scale, vulnerability patterns are becoming more and more complex, and traditional static detection tools often have a high false positive rate and false negative rate.

[0004] In recent years, with the development of machine learning and deep learning technologies, security researchers have also begun to explore how to use these technologies to detect vulnerabilities. Compared with rule-based vulnerability detection methods, machine learning and deep learning technologies have stronger adaptability and generalization ability, can effectively detect vulnerabilities, and reduce the false positive rate and false negative rate to a certain extent. Machine learning-based methods can learn potential and abstract vulnerability code patterns, and researchers use source code-based features such as header files, function calls, software complexity metrics and code changes to detect vulnerabilities. However, machine learning-based methods require experts to define vulnerability features, which is likely to lead to model bias. In contrast, deep learning-based methods do not require experts to participate in formulating program structure representation rules, and can automatically learn code patterns from source code, so as to achieve end-to-end vulnerability detection. Existing deep learning-based vulnerability code detection can be divided into two categories according to code representation, namely sequence-based methods and graph-based methods. Sequence-based representation regards code as a token sequence, and further uses models such as RNN and CNN to learn the distributed representation of code. However, source code is highly structured and hierarchical compared with natural language. Sequence-based methods can only capture the shallow syntax information of source code text, and it is difficult to learn the well-defined semantic information in the program structure. In order to better model the complex structure of code, graph-based deep learning models represent code as various graph structures, such as abstract syntax trees, control flow graphs, program dependency graphs and code property graphs, and use graph neural networks to model the structured information in code.

[0005] Although vulnerability detection methods based on machine learning and deep learning have achieved better results than traditional vulnerability detection tools, most of these methods detect code vulnerabilities at a relatively high granularity (functions, files, components, etc.) and cannot give the specific statements that may cause vulnerabilities. This makes developers still need to check a large amount of code to locate the vulnerable statements that need to be fixed. Compared with traditional rule-based vulnerability detection tools, such vulnerability detection methods lack precision. At the same time, graph-based models use message passing algorithms, and each node only communicates with its neighboring nodes. If you want to capture long-range dependencies, you need to iterate message passing multiple times. However, due to computational cost limitations, previous models usually set the number of iterations to a small number, such as less than 8. This limits the model's ability to learn long-range dependencies. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art and propose a source code vulnerability slice-level detection and statement-level positioning method based on multi-task learning. This method can overcome the above three problems in the vulnerability detection process: one is to use program slicing technology to better capture vulnerability features and improve the accuracy of vulnerability detection; the second is to model multi-task learning of whole-graph classification and node classification to accurately locate vulnerability positions at a fine granularity; the third is to build a vulnerability detection model by stacking the graph neural network GGNN and the Transformer based on the self-attention mechanism to improve the model's ability to capture long dependencies and be able to detect vulnerabilities in source code more accurately and at a finer granularity.

[0007] The technical solution of the present invention is:

[0008] A source code vulnerability slice-level detection and statement-level positioning method based on multi-task learning, and the method is as follows:

[0009] Establish a training-phase model and a detection-phase model; that is, a multi-task vulnerability detection and positioning model based on the graph neural network (GGNN) and the attention mechanism (Transformer);

[0010] Use the established training-phase model for training, and then use the detection-phase model for detection to complete source code vulnerability slice-level detection and statement-level positioning based on multi-task learning.

[0011] The input of the training-phase model is the source code file and the diff file related to the source code vulnerability repair. The diff file contains the specific statements related to the increase and decrease of the source code repair vulnerability.

[0012] The method for training the training-phase model includes:

[0013] Slice the source code file into function granularity, and extract its program dependence graph and code attribute graph for each function;

[0014] Traverse the program dependence graph, identify the nodes in the program dependence graph that contain vulnerability-sensitive elements, where the vulnerability-sensitive elements include sensitive system calls, arithmetic operations, pointer usage, array usage, and vulnerability-sensitive keywords extracted from the diff file;

[0015] Use the identified vulnerability-sensitive nodes as slicing criteria to perform forward slicing and backward slicing to obtain the sliced subgraph of the code property graph;

[0016] Use Doc2Vec to convert the source code in the nodes into vector representations, and use one-hot to encode the labels of the nodes. Concatenate the label vectors and the source code vectors as the initial vector representations of the nodes;

[0017] Segment the variable names and function names according to the camel case method and underscores, and use the BPE algorithm to compress the size of the vocabulary;

[0018] Use the node embedding matrix and edge adjacency matrix of the sliced subgraph as the input of the neural network, and use a model that combines graph neural network and Transformer for training and evaluation;

[0019] Capture the local structural information in the code through the graph neural network, capture the long-term dependence information of the code through the Transformer model, model the vulnerability detection task at the slice level as a whole graph classification task, model the vulnerability detection task at the vulnerability line level as a node classification task, and perform multi-task learning through the joint loss function.

[0020] The method for the model to perform detection in the detection stage includes:

[0021] Extract the control dependence and data dependence relationships of the target program to generate a set of sliced subgraphs;

[0022] For each set of subgraphs, generate node embeddings and input them into the trained model for detection, predicting whether each slice of the target program contains vulnerabilities and the specific vulnerability locations.

[0023] Beneficial effects

[0024] (1) A new deep learning-based source code vulnerability detection method is proposed. This method extracts the program sliced subgraphs related to vulnerability candidate elements from the entire source code based on the reachability of the program dependence graph, and performs vulnerability detection at the granularity of the sliced subgraphs, which can better capture vulnerability features and improve the detection accuracy.

[0025] (2) Model the vulnerability detection at the slice level as a whole-graph classification task, model the vulnerability location at the statement level as a node classification task, and perform multi-task learning for whole-graph classification and node classification through a joint loss function. This can not only detect vulnerabilities in code slices but also accurately locate the vulnerability positions at a fine-grained level (specific to lines).

[0026] (3) Stack the graph neural network GGNN and Transformer based on the self-attention mechanism to construct a vulnerability detection model. This model can use the message passing operation of the graph neural network to capture the local complex dependencies in the source code and use the global attention mechanism to capture the long dependencies in the source code vulnerabilities, which can greatly improve the detection ability of the model.

[0027] (4) The present invention proposes a new deep learning-based source code vulnerability detection method. This method is based on the reachability of the program dependence graph, extracts the program slice subgraphs related to vulnerability candidate elements from the entire source code, and performs vulnerability detection at the granularity of slice subgraphs, so as to better capture vulnerability features and improve the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the overall framework diagram;

[0029] Figure 2 is the vulnerability detection model;

[0030] Figure 3 is the schematic process of program slicing. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further described below with reference to the drawings and embodiments.

[0032] Embodiment

[0033] The specific model process of the present invention is as Figure 1 shown. The model is mainly divided into a training stage and a testing stage.

[0034] In the training stage, the inputs of the model are the source code file and the diff file related to the source code vulnerability repair. The diff file contains the specific statements related to the increase and decrease of the source code repair vulnerabilities. First, slice the source code file into function granularity, and for each function, extract its program dependence graph and code property graph. Then traverse the program dependence graph to identify the nodes in the program dependence graph that contain vulnerability-sensitive elements. Vulnerability-sensitive elements include sensitive system calls, arithmetic operations, pointer uses, array uses, and vulnerability-sensitive keywords extracted from the diff file. Finally, use the identified vulnerability-sensitive nodes as the slicing criterion to perform forward slicing and backward slicing to obtain the slice subgraphs of the code property graph.

[0035] Convert the source code in the nodes into vector representations using Doc2Vec, and encode the labels of the nodes using one-hot. Concatenate the label vectors and the source code vectors as the initial vector representations of the nodes. Since different code developers have different naming specifications for variables and functions during the development process, in order to accurately preserve the syntax information of the source code and prevent the OOV (Out Of Vocabulary) problem, here the variable names and function names are tokenized according to the camel case method and underscores, and the BPE algorithm is used to compress the size of the vocabulary.

[0036] Use the node embedding matrix and edge adjacency matrix of the sliced subgraphs as the input of the neural network, and use a model that combines graph neural networks and Transformers for training and evaluation. Capture the local structural information in the code through graph neural networks, capture the long-term dependency information of the code through the Transformer model, and model the vulnerability detection task at the slice level as a whole graph classification task, and model the vulnerability detection task at the vulnerability line level as a node classification task, and perform multi-task learning through a joint loss function.

[0037] In the detection stage, first extract the control and data dependency relationships of the target program to generate a set of sliced subgraphs. Then, for each group of subgraphs, generate node embeddings and input them into the trained model for detection to predict whether each slice of the target program contains vulnerabilities and the specific vulnerability locations. The implementation module of the CONFF method is as follows:

[0038] To better understand the model, use the examples described in [] to illustrate how to extract program sliced subgraphs and vectorize the nodes. The model in [] takes the code property graph as the input, but since the code property graph is too complex, for simplicity, only the program dependency graph is used in the figure. The code in the figure contains a potential integer overflow defect. It obtains a user input from the console on line 7 and may cause an overflow in the arithmetic operation on line 11.

[0039] In the program slicing stage, first generate the program dependency graph for each function, which represents the control and data dependency relationships of the program. The program dependency graph is a directed graph, where each node represents a statement or a control predicate in the source code, and a directed edge represents the control or data dependency relationship between two nodes. To better represent the syntax information of the nodes, associate each node with its node attributes, including the source code of the statement corresponding to the node and the statement type (such as expression, assignment statement, and function call). In Figure 3 (a) The second column shows the program dependency graph extracted from the example code function, where each node is concisely identified by the line number of the corresponding node statement in the source code, in and out represent the entry and exit of the program, the solid lines represent data dependency relationships, and the dashed lines represent control dependency relationships.

[0040] Extract vulnerability-sensitive elements in the program by parsing the abstract syntax tree of the program. There are five categories of vulnerability-sensitive elements: inappropriate library function calls, pointer usage, array usage, arithmetic operations, and vulnerability-sensitive keywords. As suggested in Checkmarx, some library functions are highly suspected of causing security issues. For example, memcpy and strcpy are for CWE199, and malloc and free are for CWE399. Similarly, pointer usage may lead to null pointer dereference vulnerabilities, array usage may lead to array out-of-bounds, and arithmetic operations may lead to overflow vulnerabilities.

[0041] At the same time, some user-defined functions may also pose security threats, such as OPENSSL_malloc in CVE-2014-0160, CVE-2016-2182, CVE-2016-2842, etc. Since many user-defined functions are domain-specific, such as OPENSSL_malloc and OPENSSL_free in OPENSSL, and QEMU_alloc and QEMU_free in QEMU. These code elements are obtained by identifying some representative and extensive keywords to achieve compatibility with different programs. The basic idea is to extract generalized identifiers from the Diff file, calculate the importance of the identifiers, and select vulnerability keywords according to their importance. The diff file is a file generated by the differences between the vulnerable code and the fixed code, indicating the changes required to fix the file. Therefore, the diff file is used to approximate the code statements that actually contain vulnerabilities and extract identifiers from it. Many identifiers represent the same type of object in semantics. For example, buf, buffer, dynbuf, and sigbuf represent buffer areas, while len, inp_len, and length represent the length of the object. Among them, buf and len are more general than other identifiers. By performing code-specific tokenization on the code involved in the changes, such as tokenizing according to camelCase and underscores, and then using the BPE algorithm to extract more generalized identifiers as vulnerability-sensitive keywords.

[0042] Program slicing generation

[0043] To detect vulnerabilities more accurately, the concept of candidate regions in object detection is referred to. In object detection, candidate regions refer to a series of regions that may contain objects generated on the input image in a certain way. In the classification of real-world vulnerability functions, most of the statements in a program may be statements unrelated to vulnerabilities. Therefore, it is very rough to simply classify functions, files, or components as having or not having vulnerabilities. To solve this problem, a set of subgraphs that may contain vulnerabilities is generated on the input program function through program static slicing as candidate regions for vulnerabilities. The source code function is characterized as a Code Property Graph (CPG), and a static slicing method based on the reachability of the program dependence graph is used to remove irrelevant code and graph nodes. The program slice subgraph is used as the basic granularity for detection, making the detection results more accurate.

[0044] Given the program dependence graph (PDG) of a code function, using the nodes containing vulnerability-sensitive elements as the slicing criterion, backward slicing and forward slicing are respectively performed through the breadth-first algorithm to find the code that has data dependence and control dependence relationships with the vulnerability-sensitive elements. Backward slicing traverses all the nodes and code that affect the vulnerability-sensitive elements to find the statements that may cause vulnerabilities containing the vulnerability-sensitive elements. Forward slicing traverses all the statements affected by the vulnerability-sensitive elements. Combining forward slicing and backward slicing, it can be known whether the vulnerability-sensitive elements may produce vulnerabilities and how they will affect the execution.

[0045] Finally, a slice subgraph is generated for all the statements in the function that contain vulnerability-sensitive elements. Figure 3 The third column of [figure] shows the code slice subgraph generated with the operation on line 11 as the slicing criterion. It can be seen that through program slicing, lines 5, 13, and 17 are excluded because these codes have no data dependence and control dependence relationships with line 11. At the same time, through backward slicing, it can be seen that the variable y in line 11 comes from line 7, which is user-controllable input. And through forward slicing, it can be found that the variable x in line 11 may affect the printf function in line 16, and the latter may produce a string formatting vulnerability.

[0046] Figure vectorization

[0047] The code property graph slices exist in the form of graph data. Before inputting them into the vulnerability detection neural network model, they need to be converted into vector form. The nodes in the code property graph contain two types of features. One is the label feature representing the node type, and the other is the specific code in the source code corresponding to the node. The node labels represent the specific code unit types of the nodes, such as arithmetic operations, functions, parameters, etc. There are a total of 54 different label types. For the node labels, one-hot encoding is used to represent them as 54-dimensional vectors. For the source code in the nodes, a common method is to regard the code as a sentence, decompose it into a token list, and then use word embedding techniques such as word2vec to obtain a fixed-length vector. The vocabulary mainly includes some keywords and identifiers (function names, variable names, constant values, etc.). In the real world, developers have different naming habits, and there are infinitely many possible actual variable names. Simply adding them to the vocabulary will cause the vocabulary to grow infinitely, that is, the vocabulary explosion (Vocabulary Explosion). And when encountering unknown words, the OOV (Out-Of-Vocabulary) problem will occur. This word will be regarded as an unknown word and cannot be processed, thus affecting the performance and accuracy of the model.

[0048] A common solution is symbolization. For example, SySeVR replaces user-defined function names and variable names with symbols such as FUNC and VAR. However, function names and variable names actually contain semantic information. For source code with good naming specifications, developers can simply infer their usage based on the naming of function names and variable names. To alleviate the vocabulary explosion and OOV problems, the present invention uses a subword-level expression method. According to common naming specifications, user-defined functions and variable names are divided into multiple subwords, and then the BPE (Byte Pair Encoding) algorithm is used to compress the vocabulary. Finally, the vocabulary size is compressed to 30,000. Then, a pre-trained Doc2vec model is generated using the token sequence of the code, and the pre-trained model is used to perform vector embedding on the source code in all nodes, outputting a feature matrix with dimensions of m×n, where m represents the number of nodes and n represents the dimension of the embedding vector. According to the actual experimental results, the present invention sets n to 50 dimensions. Finally, the node label vector and the source code embedding vector are concatenated to obtain the initial vector representation of the node.

[0049] Vulnerability detection model

[0050] If you want to accurately detect vulnerabilities, the model needs to be able to capture both long-range and short-range dependencies. Previous deep learning-based methods used graph neural network-based models to capture the structured information of source code. However, graph-based models use message passing algorithms, where each node only communicates with its neighboring nodes. If you want to capture long-range dependencies, multiple iterations of message passing are required. However, due to computational cost limitations, previous models usually set the number of iterations to a small number, such as less than 8. This limits the model's ability to learn long-range dependencies.

[0051] Different from graph neural network-based models, Transformer models can perform global information transfer and aggregation. There are no predefined edge relationships, but instead, they learn the relationships between nodes from scratch based on the self-attention mechanism. The vulnerability detection model of the present invention, as shown in Figure 4, takes the program slice subgraph as input, converts it into a node feature matrix and an edge adjacency matrix, and inputs them into a model group. By integrating the graph neural network model GGNN and the Transformer model, this method can effectively capture long-range and short-range dependencies in the source code and learn vulnerability pattern features. For the vulnerability detection task at the slice level, it is modeled as an entire graph classification task. The feature vector representation of the entire graph is obtained through a graph pooling layer and input into the classification layer for vulnerability detection. For the vulnerability line localization task, the node feature matrix output by the previous model is directly used, and the vulnerability line is located through a classifier. By designing a joint loss function, the model of the present invention can perform multi-task learning. Combining the advantages of GGNN and Transformer models can improve the accuracy of model vulnerability detection.

[0052] The core idea of GGNN is to iteratively update each node through a recurrent neural network (RNN) to capture the dependencies between nodes. GGNN represents the graph as a node matrix and an adjacency matrix, where the node matrix represents the initial features of each node, and the adjacency matrix represents the relationship between each node and its neighboring nodes. GGNN iteratively updates each node through a recurrent neural network, and each round of iteration includes three steps: information aggregation, transformation, and update.

[0053]

[0054]

[0055]

[0056]

[0057] Specifically, for each node v in the graph, as shown in formula (1), first initialize the node vector That is, copy the previously generated node feature vectors and pad them with 0s to a fixed length. Let T be the total number of layers of neighborhood aggregation. At each layer t (t ≤ T), all nodes perform information passing with their neighborhood nodes through edge types and directions (described by the p-th neighborhood matrix , and the number of neighborhood matrices is equal to the number of edge types). Neighborhood aggregation is shown in Equation (2), where W p is a learnable weight parameter. For each type of edge, aggregate the information of the nodes adjacent to the current node according to its adjacency matrix. Then, as shown in Equation (3), use GRU (Gated Recurrent Unit) to update the hidden state of the current node from the hidden state of the previous time step and the information of the neighborhood nodes. Here, AGG is an aggregation function, and the sum is used as the aggregation function to aggregate the information from neighbor nodes of different types of edges. Finally, as shown in Equation (4), to solve the over-smoothing problem that occurs as the depth of the graph neural network increases, the present invention refers to the JK-Net architecture to concatenate the node hidden states of each layer and obtain the final node vector representation through a linear layer.

[0058] The Transformer model used in the present invention takes the node feature vector matrix as input and can directly capture the long-range dependencies between nodes without relying on edge relationships. The Transformer model consists of multiple stacked self-attention mechanism layers, and the self-attention mechanism layer is composed of a multi-head self-attention sub-layer and a feed-forward neural network sub-layer. The calculation process of the multi-head self-attention layer is shown in Equations (5) to (7), and the calculation process of the feed-forward neural network is shown in Equation (8). There is also a residual connection and a layer normalization operation between the two sub-layers. The input of the self-attention mechanism is composed of the output of the previous layer and the position embedding vector. Specifically, the multi-head self-attention sub-layer is responsible for learning the representation of the input node sequence and learning different attention representations at different positions. The feed-forward neural network sub-layer is responsible for further transforming the output of the multi-head self-attention sub-layer to better encode the information of the input sequence. The residual connection and the layer normalization operation help to better transmit information between different layers and improve the performance of the model. The final output of the Transformer is the hidden state of each node, which contains all the information in the input nodes and can be used for vulnerability detection and vulnerability line location tasks.

[0059]

[0060]

[0061] MutiHead(Q,K,V)=Concat(head1,…,head h ) (7)

[0062] FFN(x) = max(0, xW1 + b)W2 + b2 (8)

[0063] The present invention attempts two ways to integrate the GGNN and Transformer models. One is the voting method, where the GGNN and Transformer models are trained separately, and the outputs of the models are weighted correspondingly, with the weights set to (0.5, 0.5). The other is the stacking method, with a structure such as (Transformer, GGNN, Transformer). Finally, the stacking method has slightly better results than the voting method, and the stacking structure is adopted as the final vulnerability detection model structure.

[0064]

[0065] After obtaining the appropriate source code vector representation using the model, a suitable classifier needs to be designed according to the downstream task. The goal of the present invention is to train a model that can jointly learn at both the slice level and the statement level, and they share the previous feature extraction layer based on the graph neural network and the self-attention mechanism. For the slice-level vulnerability detection task, it is modeled as a whole-graph classification task, and the global attention pooling formula (9) is used to extract the vector representation of the whole graph, where f gate (h i ) is the soft attention mechanism, which is used to determine which nodes are relevant to the current whole-graph classification task. By calculating the attention weights of all nodes and weighting them, the feature vector expression of the current graph is obtained, and then it is input into a multi-layer perceptron for vulnerability classification. For the vulnerability line localization task, the previously obtained node feature matrix is directly input into a multi-layer perceptron for detection at the vulnerability line level. If a code slice subgraph is predicted to be non-vulnerable, it should not contain vulnerable statements. The slice-level output (0 represents no vulnerability, 1 represents vulnerability) is element-wise multiplied with the vulnerability line localization result to balance the conflict between the slice-level and statement-level detection results.

[0066] The present invention proposes a multi-task model that can simultaneously perform slice-level vulnerability detection and vulnerability line localization tasks. For these two tasks, the present invention adopts the cross-entropy loss function. Since the difference between vulnerable code and secure code is usually very small, the present invention introduces a contrastive loss function to maximize the distance between vulnerable samples and non-vulnerable samples. It tries to make the similarity between positive samples and negative samples as close as possible during training, while making the similarity between positive samples and all negative samples as far as possible, so as to enhance the discrimination ability of the classifier. The loss function is designed as shown in formulas (10)-(13). Where L f is the loss function for the vulnerability detection task, and L s is the loss function for the vulnerability line localization task. Due to the data imbalance in the real world, w1 and w2 are used for corresponding weighting, Lc is the contrast loss function. For a given sample, the y in the same batch with the same label as it j is regarded as a positive example, and the y with a different label from it k is regarded as a negative example. The purpose of this contrast loss function is to map samples of the same category to a similar embedding space and map samples of different categories to different embedding spaces. The final combined loss function L is the weighted sum of these three loss functions.

[0067]

[0068]

[0069]

[0070] L = L f + L s + L c (13)

[0071] To achieve slice-level vulnerability detection, the present invention models it as a whole-graph classification task and models statement-level vulnerability localization as a node classification task. Through multi-task learning of whole-graph classification and node classification using a combined loss function, vulnerabilities can be detected on code slices and the vulnerability locations can be accurately located at a fine-grained level (specific to lines).

[0072] To build a vulnerability detection model, as Figure 2 shown, the present invention stacks the graph neural network GGNN and the Transformer based on the self-attention mechanism. This model can utilize the message passing operation of the graph neural network to capture local complex dependencies in the source code and utilize the global attention mechanism to capture long dependencies in source code vulnerabilities, thereby greatly improving the detection ability of the model.

[0073] In summary, the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for slice-level detection and statement-level localization of source code vulnerabilities based on multi-task learning, characterized in that The method is as follows: Build a training-phase model and a detection-phase model; Use the built training-phase model for training, and then use the detection-phase model for detection to complete the slice-level detection and statement-level location of source code vulnerability based on multi-task learning.

2. The method for slice-level detection and statement-level localization of source code vulnerabilities based on multi-task learning according to claim 1, characterized in that: The input of the training-phase model is the source code file and the diff file related to the source code vulnerability repair. The diff file contains the specific statements related to the increase and decrease of the source code repair vulnerability.

3. The method for slice-level detection and statement-level localization of source code vulnerabilities based on multi-task learning according to claim 1 or 2, characterized in that: The method for training the training-phase model includes: Slice the source code file into function granularity, and extract its program dependence graph and code property graph for each function; Traverse the program dependence graph to identify the nodes containing vulnerability-sensitive elements in the program dependence graph. The vulnerability-sensitive elements include sensitive system calls, arithmetic operations, pointer usage, array usage, and vulnerability-sensitive keywords extracted from the diff file; Use the identified vulnerability-sensitive nodes as slicing criteria to perform forward slicing and backward slicing to obtain the sliced subgraph of the code property graph; Use Doc2Vec to convert the source code in the nodes into vector representations, and use one-hot to encode the labels of the nodes. Concatenate the label vectors and the source code vectors as the initial vector representations of the nodes; Segment the variable names and function names according to the camel case method and underscores, and use the BPE algorithm to compress the vocabulary size; Use the node embedding matrix and edge adjacency matrix of the sliced subgraph as the input of the neural network, and use a model combining graph neural network and Transformer for training and evaluation; Capture the local structured information in the code through the graph neural network, capture the long-term dependence information of the code through the Transformer model, model the slice-level vulnerability detection task as a whole-graph classification task, model the vulnerability detection task at the vulnerability line level as a node classification task, and perform multi-task learning through the joint loss function.

4. The method for slice-level detection and statement-level localization of source code vulnerabilities based on multi-task learning according to claim 1, characterized in that: The method for the detection-phase model to perform detection includes: Extract the control dependence and data dependence relationships of the target program to generate a set of sliced subgraphs; For each group of subgraphs, generate node embeddings and input them into the trained model for detection to predict whether each slice of the target program contains vulnerabilities and the specific vulnerability locations.

Citation Information

Cited By

  • Vulnerability mining and dynamic updating method and device based on machine learning and medium

    CN121681322A

  • Machine learning-based vulnerability mining and dynamic updating method, device and medium

    CN121681322B

  • Code vulnerability positioning method and system based on cooperation of coarse and fine granularities

    CN122241724A