Source code-oriented fine-grained vulnerability detection method and system
By modeling the vulnerability detection task as a whole graph classification task and employing a graph neural network with an ordered message passing mechanism, the accuracy and efficiency issues of fine-grained vulnerability detection in existing technologies are solved, achieving efficient vulnerability localization and detection.
Patent Information
- Application Number
- CN202510932086.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-28
AI Technical Summary
Existing deep learning-based vulnerability detection technologies struggle to achieve fine-grained vulnerability detection, failing to pinpoint specific vulnerable statements. Furthermore, traditional slicing methods introduce redundancy and noise, impacting detection accuracy and efficiency.
The vulnerability detection task at the slice level is modeled as a whole graph classification task. A graph neural network model with an ordered message passing mechanism is used, combined with static analysis techniques to slice the program, generating more accurate slice results. Long-distance dependencies are captured by Ordered GNN.
It achieves high-accuracy, fine-grained vulnerability detection, accurately locates vulnerability positions, significantly improves detection efficiency and accuracy, and reduces interference from redundant information.
Smart Images

Figure CN120850294A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vulnerability detection, specifically relating to a fine-grained vulnerability detection method and system for source code. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of information technology, software systems are increasingly widely used in finance, healthcare, national defense, transportation, and other fields. Information security has also become a major concern, with frequent cybersecurity incidents such as hacker attacks, user privacy breaches, and theft of information and assets posing significant threats to individual privacy, corporate interests, and even national security. The root cause of these information security problems lies in security vulnerabilities within software systems.
[0004] Software vulnerabilities refer to operational anomalies caused at various stages of software development due to intentional or unintentional design flaws, non-standard coding practices, etc., by developers. These vulnerabilities may stem from logical errors in the code, improper memory management, insufficient input validation, and other issues. With the vigorous development of the open-source community, the speed and scope of software vulnerability propagation have further expanded. The widespread use of open-source software means that the impact of vulnerabilities is no longer limited to a single software or system, but can potentially affect the entire ecosystem.
[0005] Therefore, how to efficiently and accurately detect vulnerabilities in source code has become an important research direction in the field of information security. In recent years, deep learning technology has made tremendous progress. Compared with traditional rule-based vulnerability detection technology, deep learning technology has stronger automation and adaptability, and its vulnerability detection accuracy is higher. It can automatically extract potential vulnerability features in code, learn vulnerability patterns automatically through training, and learn complex vulnerability patterns from large amounts of data, enabling it to adapt to new vulnerabilities and cope with unknown attacks.
[0006] Existing deep learning-based vulnerability detection technologies can be broadly categorized into two types: text sequence-based and graph-based vulnerability detection technologies. In its early stages, deep learning-based methods could only perform vulnerability detection at a high granularity (such as files, functions, and code blocks), unable to provide finer-grained results. While this coarse-grained detection approach could quickly locate potentially vulnerable code regions, its large granularity often prevented it from pinpointing specific vulnerable statements, requiring developers to perform further manual analysis during vulnerability remediation, increasing both cost and time.
[0007] As research progresses, an increasing number of methods are focusing on fine-grained vulnerability detection. These methods can further pinpoint specific vulnerable statements based on the results of coarse-grained vulnerability detection, thus significantly improving the accuracy and practicality of vulnerability detection. However, the semantic and structural information of a single statement is often insufficient to accurately identify vulnerable statements. Therefore, current research frequently uses program slicing to obtain the code context for vulnerability detection. Traditional slicing methods focus on lines of code that may affect the target, which can result in large and redundant slices, introducing significant noise and complex statement dependencies. This limits the accuracy of existing methods in fine-grained vulnerability detection. Summary of the Invention
[0008] To address the aforementioned problems, this invention proposes a fine-grained vulnerability detection method and system for source code. This invention models the slice-level vulnerability detection task as a whole-graph classification task and the vulnerability statement location task as a node classification task. This invention achieves high accuracy in vulnerability detection and can detect and precisely locate vulnerabilities at a fine-grained level.
[0009] According to some embodiments, the present invention adopts the following technical solution: A fine-grained vulnerability detection method for source code includes the following steps: Obtain existing program source code and vulnerability-sensitive elements as training data; The program source code in the training data is compiled into an intermediate representation. Based on the central node containing vulnerability-sensitive elements, the code value flow graph is sliced inter-process to extract vulnerability-related information from the code. The generated slice subgraphs are represented by vectors, transforming the graph structure information into a format that can be input into the vulnerability detection model; Train the vulnerability detection model using the processed training data; Extract the slice subgraph of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
[0010] As an alternative implementation, pointer analysis is performed on the formed intermediate representation, and it is converted into a static single-assignment form to generate the value flow graph of the program.
[0011] As an alternative implementation, the process of performing inter-process slicing on the code value flow graph based on the central node containing the vulnerability-sensitive element includes: locating the line of code containing vulnerability-sensitive information in the program source code, and then locating the node in the value flow graph; using the node containing the vulnerability-sensitive element as the slice center, performing backward slicing and forward slicing respectively using a breadth-first search algorithm to find the code that has a value flow dependency relationship with the vulnerability-sensitive element.
[0012] As a further implementation, backward slicing traverses all nodes that affect vulnerability-sensitive elements to find statements that may cause vulnerabilities in nodes containing vulnerability-sensitive elements; forward slicing traverses all statements affected by vulnerability-sensitive elements.
[0013] As an alternative implementation, the process of vectorizing the generated slice subgraph includes: after the static analysis slice is completed, obtaining the instruction slices related to the vulnerability; by adding debugging information during the compilation process, obtaining the line numbers of the source code corresponding to these instruction-level slices; converting the source code into a code attribute graph; by matching the line numbers obtained from the slices with the elements in the code attribute graph, obtaining the slice subgraph of the code attribute graph; converting the slice subgraph into vector form, where each node in the code attribute graph represents a statement or code element; for each node in the graph, generating the node embedding of its corresponding statement; and combining the edge adjacency matrix of the graph as the input to the model.
[0014] As an alternative implementation, the vulnerability detection model is based on a graph neural network model, which sorts the messages passed to the node representation. For a central node v, the neighboring nodes within a certain number of hops are represented as a rooted tree, denoted as v. The (k-1)th order root tree of the central node v is a subtree of the k-th order root tree, with the following relationship:
[0015] The vector representation within K hops of the corresponding node v in a K-order rooted tree, in the vector representation of node v. In the process, the neurons encoding the k-th order root tree include the neurons encoding the (k-1)-th order root tree; The final embedding of node v in a K-layer graph neural network model Divided into K+1 blocks by K+1 dividing points, where the first Each neuron encodes information about node v itself. for The dividing point; Define the gate vector To control the encoding of neighbor node information into the neurons of the current node, using This is used to record the specific location where neighbor information is embedded into the current node, thus sequentially encoding the neighbor information of each hop into the final representation vector. middle.
[0016] As a further implementation, the expectation vector is used. Represented as each position in the node embedding as The cumulative sum of probabilities is used to ensure the differentiability of the entire model:
[0017] in Indicates the cumulative sum from right to left. softmax Composite operators, This indicates the fusion of two vectors. and The function represents a linear projection with bias. Represent two vectors and Serial; Then, the information of the current node and its neighboring nodes is aggregated by dividing the data into blocks according to the distance from each neighboring node to the current node.
[0018] As an alternative implementation, for slice-level vulnerability detection tasks, it is modeled as a whole-graph classification task. The feature vector representation of the whole graph is obtained through the graph pooling layer and input into the classification layer for vulnerability detection. For vulnerability row localization tasks, the node feature matrix output by the model is used to locate the vulnerability row through the classifier.
[0019] A fine-grained vulnerability detection system for source code, comprising: The data acquisition module is configured to acquire existing program source code and vulnerability-sensitive elements as training data. The slicing module is configured to compile the program source code in the training data into an intermediate representation, and perform inter-process slicing on the code value flow graph based on the central node containing vulnerability-sensitive elements to extract vulnerability-related information from the code. The conversion module is configured to perform vector representation on the generated slice subgraphs, converting the graph structure information into a format that can be input into the vulnerability detection model; The training module is configured to train the vulnerability detection model using processed training data; The vulnerability detection module is configured to extract slice subgraphs of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
[0020] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the method described above.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention models slice-level vulnerability detection tasks as whole-image classification tasks and vulnerability statement location tasks as node classification tasks. This invention achieves high accuracy in vulnerability detection and can detect and precisely locate vulnerabilities at a fine-grained level, down to the line level.
[0022] This invention optimizes the slicing algorithm to effectively remove redundant information from program code, thereby significantly improving the efficiency and accuracy of vulnerability detection. The method combines static analysis techniques to analyze the "definition-usage" relationships of variables in the program, generating more accurate slicing results. This reduces the introduction of redundant code, allowing the model to focus more on key code segments related to vulnerabilities, thus improving detection accuracy.
[0023] To address the oversmoothing problem inherent in traditional graph neural networks when handling long-distance dependencies between non-adjacent nodes, this invention introduces an ordered message passing mechanism into the vulnerability detection model, enabling more accurate node classification. Based on this ordered message passing mechanism, the vulnerability detection model aligns the hierarchical structure of the central node's root tree with the ordered neurons in the node representation, embedding the transmitted information into the nodes in an orderly manner according to their distance from the nodes. This model's results are easier to interpret and exhibit stronger performance. This mechanism not only improves the model's ability to understand complex code structures but also further enhances the accuracy of vulnerability detection.
[0024] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. Attached Figure Description
[0025] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0026] Figure 1 This is a schematic diagram of a fine-grained vulnerability detection method for source code in one embodiment; Figure 2 This is a schematic diagram of an ordered message passing mechanism according to one embodiment; Figure 3 This is a schematic diagram of a vulnerability detection model for one embodiment. Detailed Implementation
[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0029] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0030] Where there is no conflict, the embodiments and features described in this application may be combined with each other.
[0031] Example 1 As described in the background section, current methods using program slicing still cannot effectively remove redundant information, and irrelevant information can interfere with vulnerability detection results. The semantic and structural information of a single statement is often difficult to use directly for vulnerability detection. This is because vulnerabilities typically involve complex dependencies between multiple statements, and analyzing a single statement alone is unlikely to capture the complete vulnerability context. Therefore, current research often employs program slicing techniques to obtain the context of vulnerability detection code. Program slicing techniques extract code fragments related to the target statement by analyzing the data flow and control flow in the code, thus providing richer contextual information for vulnerability detection. However, traditional program slicing techniques focus on a set of code lines that affect the target, which may produce large redundant slices and introduce a lot of noise when capturing dependencies between statements. This redundancy and noise not only increases computational overhead but may also reduce the accuracy of vulnerability detection. For example, traditional slicing techniques may include code fragments unrelated to the vulnerability, causing interference to the model during learning and thus affecting the accuracy of the detection results.
[0032] On the other hand, traditional graph neural networks suffer from oversmoothing, which limits the accuracy of statement-level vulnerability localization. Graph neural networks (GNNs) have limitations in handling long-distance relationships between non-adjacent nodes. Although global information of the graph (i.e., long-range node dependencies) can be captured by stacking GNN layers, deep GNNs are prone to oversmoothing, causing node embeddings with different features to become similar, thus obfuscating node differences. This oversmoothing problem is particularly prominent in vulnerability detection because vulnerabilities often involve complex dependencies between multiple statements, which may be distributed across different parts of the code. If GNNs cannot effectively capture these long-distance dependencies, the model will perform poorly in vulnerability detection. To address this issue, some studies have proposed improved GNN architectures, such as introducing attention mechanisms to enhance the model's ability to capture long-distance dependencies. In addition, some studies have attempted to combine GNNs with other techniques (such as symbolic execution or program slicing) to compensate for the shortcomings of GNNs in handling complex code structures.
[0033] To address the aforementioned issues, this embodiment provides a fine-grained vulnerability detection method oriented towards source code, comprising the following steps: Obtain existing program source code and vulnerability-sensitive elements as training data; The program source code in the training data is compiled into an intermediate representation. Based on the central node containing vulnerability-sensitive elements, the code value flow graph is sliced inter-process to extract vulnerability-related information from the code. The generated slice subgraphs are represented by vectors, transforming the graph structure information into a format that can be input into the vulnerability detection model; Train the vulnerability detection model using the processed training data; Extract the slice subgraph of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
[0034] The following is a detailed introduction: (e.g.) Figure 1 As shown, it mainly includes two phases: the training phase and the detection phase. The training phase takes a set of program source code and vulnerability-sensitive elements as input. The training phase includes the following steps: (1) Compile: Compile the source code of the training data into an intermediate representation to facilitate subsequent operations; (2) Slice the program: Using the central node containing vulnerability-sensitive elements as a reference, slice the code value flow graph inter-process to extract vulnerability-related information from the code; (3) Generate vector representations of the sliced subgraphs: Transform the graph structure information into a format that can be input into the model; (4) Train the model using the processed training data.
[0035] During the detection phase, slice subgraphs of test data are extracted, node embeddings are generated, and the data are input into a trained model for detection to predict whether the slices of the target program contain vulnerabilities and their specific locations.
[0036] In a software system, vulnerabilities are often closely related to syntactic features such as pointer usage, array usage, library function calls, and arithmetic expressions. These syntactic features are key points in vulnerability detection because they are often high-risk areas for vulnerabilities.
[0037] To more comprehensively capture these potential vulnerability patterns, this invention uses the aforementioned vulnerability syntax features to guide vulnerability detection and analysis. First, the source code is compiled into LLVM IR, a prerequisite for subsequent analysis, and then the SVF framework is applied for subsequent static analysis. To generate a more accurate program value flow graph, this embodiment performs pointer analysis on the IR and converts it into a static single-assignment form. Based on the pointer analysis, the program's value flow graph is generated. The value flow graph is a directed graph where each node represents an instruction, function entry / exit point, parameter information at the call point, and each edge represents the value flow dependency from variable definition to its usage.
[0038] The slicing strategy of this embodiment is given below. Given a program P and a target point t, our slicing algorithm computes a subgraph containing the data flow related to the target point. Assume the control flow graph of the program is represented as follows: , It is a collection of program nodes. Value flow graph represents the control dependencies between nodes. It has the same nodes as the control flow graph, but it includes data dependencies instead of control dependencies. We define data dependency edges based on a compact slicing strategy: exist definition x in Used at Not in and The nodes between are defined in and Here, x is a node in the code graph, and 'x' is a variable in program P. To perform accurate slicing, it's necessary to know which variables each pointer points to; without this information, it's impossible to determine which variables are defined and used at each program point. Therefore, this embodiment performs flow-insensitive and context-insensitive pointer analysis to derive data dependencies. (Obtaining the graph...) Next, we slice the graph to obtain sliced subgraphs containing nodes related to the target location:
[0039] The slicing algorithm used in this embodiment first locates the lines of code containing vulnerability-sensitive information in the source code. These lines of code then pinpoint nodes in the value flow graph. Using the nodes containing the vulnerability-sensitive elements as slice centers, a breadth-first search algorithm is used to perform backward and forward slicing to find code with value flow dependencies on the vulnerability-sensitive elements. Backward slicing traverses all nodes affecting the vulnerability-sensitive elements, finding statements that might cause vulnerabilities in nodes containing the vulnerability-sensitive elements. Forward slicing traverses all statements affected by the vulnerability-sensitive elements. Combining forward and backward slicing, it can be determined whether the vulnerability-sensitive element might cause a vulnerability and how it will affect execution. Given a given scenario where all possible functions have already been identified through pointer analysis, this slicing process is cross-function, thus obtaining a complete vulnerability-related slice across functions.
[0040] After static analysis of the slices is completed, the instruction slices related to the vulnerability in the IR can be obtained. By adding debugging information during the compilation process, the line numbers of the source code corresponding to these instruction-level slices can be obtained.
[0041] Adding the "-g" option to the compiler directive generates debugging information (such as variable names, function names, and source code line numbers) in the intermediate representation generated during compilation. This allows for direct analysis of the intermediate representation to pinpoint problems in the source code.
[0042] Finally, the source code is transformed into a code attribute graph. By mapping the line numbers obtained from the slices to the elements in the code attribute graph, a slice subgraph of the code attribute graph is obtained. The code attribute graph slices exist in graph data form, which needs to be converted into vector form before being input into the vulnerability detection neural network model. Nodes in the code attribute graph represent a statement or code element (operator, variable, etc.). For each node in the graph, the node embedding of its corresponding statement is generated, and combined with the graph's edge adjacency matrix, it serves as the input to the model.
[0043] Ordered message passing mechanism: Ordered GNN is a graph neural network based on an ordered message passing mechanism. It orders the messages passed to the node representation and uses specific neuron blocks for messages passed within a specific hop. This is achieved by aligning the hierarchical structure of the root tree of the central node with the ordered neuron blocks of its node representation. During each round of message passing, the representation of the central node is updated according to the level of its neighboring nodes, and the information of the neighboring nodes is encoded sequentially in the embedding of the central node.
[0044] like Figure 2 As shown, for a central node v, its neighboring nodes within K hops can be represented as a rooted tree, denoted as v. Clearly, the (k-1)th order root tree of v is a subtree of the k-th order root tree, with the following relationship:
[0045] The vector representation within K hops of the corresponding node v in a K-order rooted tree, in the vector representation of node v. In this context, the neurons encoding the k-th order root tree must include neurons encoding the (k-1)-th order root tree. Specifically, assume that the vector encoding v contains... For each neuron, the information of the (k-1)th order root tree of node v will be encoded into the preceding neurons. In each neuron, for the root tree at the next level, the previous neurons will be encoded. In each neuron, due to yes The subtree, therefore Therefore, in the final node representation of v, the previous... Each neuron encodes information about its first k-1 hop neighbors. Each neuron encodes information about its first k hop neighbors at the two split points. and The neurons between them encode information about the k-th hop neighbor.
[0046] Therefore, the final embedding of node v in the K-layer GNN It will be divided into K+1 blocks by K+1 dividing points, of which the first Each neuron encodes information about node v itself. for The dividing point, and:
[0047] Dividing point It is an index that will divide the D-dimensional node embedding into two parts, [0, The neurons in the first k layers encode information about their first k neighbors. Define a D-dimensional gating vector, where the first k neighbors are... The first element is 1, followed by 0, which means selecting the neurons to be encoded in the first k layers:
[0048] The information of the (k-1)th layer Preserve the embedding at the k-th layer Before In each neuron, and The later neurons encode newly aggregated neighbor information, thus separating the information from each hop. Then, when moving to the next layer... The information will be encoded into the first D neurons. In each neuron, then actually arrive The neurons between them actually only contain The information, specifically the information of the k-th hop, is used to sequentially encode the neighbor information of each hop into the final representation vector. middle.
[0049] Fine-grained vulnerability detection model To accurately detect vulnerabilities, models need the ability to precisely capture the features of vulnerable statements. Previous deep learning-based methods used graph neural network-based models to capture the structured information of source code. However, graph-based models use message passing algorithms, where each node only communicates with its neighbors. To capture long-range dependencies, multiple iterations of message passing are required. As the model depth increases, these iterations can lead to nodes of different categories having the same embeddings, causing confusion between different node features. This limits the model's ability to capture the features of vulnerable statements, resulting in poor performance in node classification tasks for statement-level vulnerability detection.
[0050] The vulnerability detection model in this embodiment is as follows: Figure 3 As shown, the program slice subgraph is used as input, which is transformed into a node feature matrix and an edge adjacency matrix and then input into the model group. After passing through the input layer and normalization processing, it is fed into the subsequent OrderedGNN module. Normalization can make different features have similar scales, avoid gradient vanishing and gradient exploding problems, make network training more stable, accelerate model convergence, prevent overfitting, and improve the model's generalization ability.
[0051] In addition to the characteristics of the nodes, this embodiment distinguishes different edges by updating the nodes and their neighboring nodes under each edge type.
[0052] Therefore, the model in this embodiment is more suitable for heterogeneous graphs. The method in this embodiment integrates Ordered GNN, a neural network that implements an ordered message passing mechanism. When capturing long-distance dependencies, the features between nodes can be effectively distinguished, thus avoiding the need for multiple layers of network transmission while still maintaining high model performance.
[0053] This embodiment uses a binary gating vector. To control the encoding of neighbor node information into the neurons of the current node, using This is used to record the specific location where information from neighboring nodes is embedded into the current node. Because of vectors... Because it originates from discrete operations, its learning is non-differentiable, so the expected vector is used. Represented as each position in the node embedding as The cumulative sum of probabilities is used to ensure the differentiability of the entire model.
[0054]
[0055] in Indicates the cumulative sum from right to left. softmax Composite operators, This indicates the fusion of two vectors. and The function represents a linear projection with bias. Represent two vectors and Serial connection. Due to It is learnable. To ensure that neighbor information within different layers can be embedded into the current node vector in an orderly manner, it uses... The operator is guaranteed to be differentiable.
[0056]
[0057] Then, the information of the current node and its neighboring nodes is aggregated by dividing the data into blocks according to the distance from each neighboring node to the current node: . in Represents a learnable gated vector. The information refers to the neighbor nodes obtained above. We add residual connections between networks in each layer of the Ordered GNN. After adding residual connections, the output of one layer in the network is not only passed to the next layer, but also directly connected to deeper layers. This allows deeper layers to directly obtain the original information, or "residual information," during updates, enabling the model to learn more feature combinations. For slice-level vulnerability detection tasks, we model it as a whole-graph classification task. We obtain the feature vector representation of the whole graph through a graph pooling layer and input it into the classification layer for vulnerability detection. For vulnerability row localization tasks, we directly use the node feature matrix output by the previous model and perform vulnerability row localization through a classifier.
[0058] Given the superior performance and dominant position of MLPs in classifiers, this embodiment further utilizes MLPs to train the classifier. The MLP used in this embodiment consists of three layers: an input layer, an intermediate layer, and an output layer. Here, the output of the Ordered GNN module is used as the input to the MLP layer, and the transformation process of the node representation can be expressed as: . in and These are learnable parameters. It is a vector representation of a node. This refers to the number of layers in the Ordered GNN. To maximize the distance between vulnerable and non-vulnerable samples, this embodiment uses the cross-entropy loss function to measure the difference between the model's predicted probability distribution and the true labels:
[0059] in It is the number of nodes. It is the number of categories. It is the one-hot encoding of the actual label of node i. This represents the probability predicted by the model that node i belongs to node j. Finally, the output layer will use the graph features output from the intermediate layers for prediction:
[0060] in and These are all learnable parameters.
[0061] In summary, this embodiment significantly improves the accuracy and efficiency of vulnerability detection by combining program slicing technology and ordered message passing mechanisms. One of the innovations of this embodiment is its ability to effectively capture complex dependencies in code and reduce interference from irrelevant information. This embodiment models subgraph-level vulnerability detection as a graph classification task and vulnerability localization as a graph node classification task, enabling more precise location of vulnerable statements, thus achieving excellent performance in both subgraph-level vulnerability detection and statement-level vulnerability localization tasks.
[0062] This embodiment removes redundant code, focusing on critical code segments relevant to the vulnerability. This invention combines pointer analysis technology, analyzing the "definition-usage" relationships of variables in the code, and applying reachability analysis algorithms to obtain program slices based on the program's control flow and data flow. The slicing method of this invention uses four types of vulnerability characteristics—library function / API calls, array usage, pointer usage, and arithmetic expressions—as slicing criteria, enabling it to distinguish between different vulnerability types for inter-procedural slicing. Experimental verification shows that the slicing method of this invention effectively retains vulnerability-related information and removes vulnerability-irrelevant code. This method significantly reduces the introduction of noise, improving detection efficiency and accuracy.
[0063] This embodiment aligns the root tree hierarchy corresponding to the central node with the ordered neurons in the node representation, embedding them into the nodes in an ordered manner according to the distance of the embedded information. This enables the model to better capture the dependencies between distant nodes and distinguish the features of different types of nodes. Experimental results show that the model of this invention has advantages in distinguishing the features of different node types and outperforms existing vulnerability detection methods in node classification tasks based on vulnerability statement location.
[0064] Example 2 A fine-grained vulnerability detection system for source code, comprising: The data acquisition module is configured to acquire existing program source code and vulnerability-sensitive elements as training data. The slicing module is configured to compile the program source code in the training data into an intermediate representation, and perform inter-process slicing on the code value flow graph based on the central node containing vulnerability-sensitive elements to extract vulnerability-related information from the code. The conversion module is configured to perform vector representation on the generated slice subgraphs, converting the graph structure information into a format that can be input into the vulnerability detection model; The training module is configured to train the vulnerability detection model using processed training data; The vulnerability detection module is configured to extract slice subgraphs of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
[0065] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROM It takes the form of a computer program product implemented on (such as optical memory, etc.).
[0066] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0068] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A fine-grained vulnerability detection method for source code, characterized in that, Includes the following steps: Obtain existing program source code and vulnerability-sensitive elements as training data; The program source code in the training data is compiled into an intermediate representation. Based on the central node containing vulnerability-sensitive elements, the code value flow graph is sliced inter-process to extract vulnerability-related information from the code. The generated slice subgraphs are represented by vectors, transforming the graph structure information into a format that can be input into the vulnerability detection model; Train the vulnerability detection model using the processed training data; Extract the slice subgraph of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
2. The fine-grained vulnerability detection method for source code as described in claim 1, characterized in that, The intermediate representation is analyzed by pointers and converted into a static single-assignment form to generate the value flow graph of the program.
3. The fine-grained vulnerability detection method for source code as described in claim 1, characterized in that, The process of performing inter-process slicing on the code value flow graph, based on the central node containing the vulnerability-sensitive element, includes: locating the line of code containing vulnerability-sensitive information in the program source code, and then locating the node in the value flow graph; using the node containing the vulnerability-sensitive element as the slice center, performing backward slicing and forward slicing respectively using a breadth-first search algorithm to find code that has value flow dependencies on the vulnerability-sensitive element.
4. The fine-grained vulnerability detection method for source code as described in claim 3, characterized in that, backward... The slice iterates through all nodes that affect the vulnerability-sensitive element, finding statements that may cause vulnerabilities in nodes containing the vulnerability-sensitive element; the forward slice iterates through all statements affected by the vulnerability-sensitive element.
5. The fine-grained vulnerability detection method for source code as described in claim 1, characterized in that, The process of vectorizing the generated slice subgraphs includes: after static analysis of the slices, obtaining instruction slices related to the vulnerabilities; adding debugging information during the compilation process to obtain the line numbers of the source code corresponding to these instruction-level slices; converting the source code into a code attribute graph; obtaining slice subgraphs of the code attribute graph by mapping the line numbers obtained from the slices to the elements in the code attribute graph; converting the slice subgraphs into vector form; where each node in the code attribute graph represents a statement or code element; generating the node embedding of its corresponding statement for each node in the graph; and combining the edge adjacency matrix of the graph as the input to the model.
6. The fine-grained vulnerability detection method for source code as described in claim 1, characterized in that, The vulnerability detection model is based on a graph neural network model. It sorts the messages passed to the node representations. For a central node v, the neighboring nodes within a certain number of hops are represented as a rooted tree, denoted as vv. The (k-1)th order root tree of the central node v is a subtree of the k-th order root tree, with the following relationship: The vector representation within K hops of the corresponding node v in a K-order rooted tree, in the vector representation of node v. In the process, the neurons encoding the k-th order root tree include the neurons encoding the (k-1)-th order root tree; The final embedding of node v in a K-layer graph neural network model Divided into K+1 blocks by K+1 dividing points, where the first Each neuron encodes information about node v itself. for The dividing point; Define the gate vector To control the encoding of neighbor node information into the neurons of the current node, using This is used to record the specific location where neighbor information is embedded into the current node, thus sequentially encoding the neighbor information of each hop into the final representation vector. middle.
7. A fine-grained vulnerability detection method for source code as described in claim 6, characterized in that, Using the expected vector Represented as each position in the node embedding as The cumulative sum of probabilities is used to ensure the differentiability of the entire model: in Indicates the cumulative sum from right to left. softmax Composite operators, This indicates the fusion of two vectors. and The function represents a linear projection with bias. Represent two vectors and Serial; Then, the information of the current node and its neighboring nodes is aggregated by dividing the data into blocks according to the distance from each neighboring node to the current node.
8. The fine-grained vulnerability detection method for source code as described in claim 1, characterized in that, For slice-level vulnerability detection tasks, they are modeled as whole-graph classification tasks. The feature vector representation of the whole graph is obtained through the graph pooling layer and input into the classification layer for vulnerability detection. For vulnerability row localization tasks, the node feature matrix output by the model is used to locate the vulnerability row through the classifier.
9. A fine-grained vulnerability detection system for source code, characterized in that, include: The data acquisition module is configured to acquire existing program source code and vulnerability-sensitive elements as training data. The slicing module is configured to compile the program source code in the training data into an intermediate representation, and perform inter-process slicing on the code value flow graph based on the central node containing vulnerability-sensitive elements to extract vulnerability-related information from the code. The conversion module is configured to perform vector representation on the generated slice subgraphs, converting the graph structure information into a format that can be input into the vulnerability detection model; The training module is configured to train the vulnerability detection model using processed training data; The vulnerability detection module is configured to extract slice subgraphs of the target program's source code, generate node embeddings, and input them into a trained vulnerability detection model for detection, predicting whether the slice of the target program contains vulnerabilities and their specific locations.
10. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the steps of the method according to any one of claims 1-8.