Software supply chain risk monitoring and early warning method and device
By extracting graph structure information and computing structure features of the source code of the open source component that the software depends on, and detecting it in combination with the graph convolutional neural network model, the problem of low accuracy of code vulnerability detection in the existing technology is solved, and higher detection comprehensiveness and accuracy are achieved.
Patent Information
- Application Number
- CN202510422421.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has low accuracy in code vulnerability detection and cannot effectively capture the syntax structure and semantic information of the code, resulting in a high rate of missed detection and false detection.
By extracting the graph structure information from the open source component source code that the software depends on, building a composite graph, and computing the structural feature value of each node. Combined with the pre-trained embedding representation model, node information is vectorized, graph embedding matrix and structural feature vector are generated, and graph convolutional neural network model is input for detection.
Through the extraction of graph structure information and the calculation of structural features, the syntax structure and semantic information of the code can be better preserved, the comprehensiveness and accuracy of vulnerability detection can be improved, and the missed detection and false detection rates can be reduced.
Smart Images

Figure CN119939607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vulnerability detection, and in particular to a software supply chain risk monitoring and early warning method and device. Background Art
[0002] The open source software supply chain refers to the supply relationship network formed by dependencies, combinations, etc., including upstream communities, source code packages, binary packages, third-party component distribution markets, application software distribution markets, and developers and maintainers, communities, foundations, etc. involved in the development and operation of open source software. Software supply chain security includes two parts: one is the existence of vulnerabilities in the introduced open source component packages, and the other is the existence of malicious code carefully crafted by attackers in the introduced open source component packages. Third-party component dependencies and code reuse may cause vulnerabilities in upstream open source components to be introduced into downstream components and application software. For example, copying code from GitHub and pulling open source component packages from PyPI and Maven may introduce vulnerabilities.
[0003] Patent CN116226864A discloses a code vulnerability detection method and system for network security, which converts code text into vectors and inputs them into a machine learning model to obtain detection results. However, the code has complex syntax and semantics, and simple vectorization fails to fully reflect the code's grammatical structure (such as loops, conditional statements) and semantic information (such as variable scope, data dependency). This makes the model's missed detection and false detection rates high, that is, the existing technology has the problem of low accuracy. Summary of the invention
[0004] The purpose of the present invention is to solve the problem of low accuracy mentioned in the above background technology, and to propose a software supply chain risk monitoring and early warning method and device.
[0005] A first aspect of the present invention provides a software supply chain risk monitoring and early warning method, the method comprising: Obtaining the source code of the open source component that the software depends on, and preprocessing the target code fragment to obtain the target source code; the target code fragment is any function segment in the source code of the open source component; Extracting graph structure information from the target source code to obtain a composite graph; According to the edges of the composite graph, a structural characteristic value of each node in the composite graph is calculated; According to the pre-trained embedding representation model, the node information of the target node is vectorized to obtain a node feature vector of the target node; the target node is any node in the composite graph; Obtaining a graph embedding matrix and a structural eigenvector according to the node eigenvectors and structural eigenvalues of all nodes in the composite graph; The graph embedding matrix and the structural feature vector are used as inputs of a preset detection model to obtain a detection result; the detection result includes code vulnerabilities and code security; If the detection result shows that there is a vulnerability in the code, an early warning message is generated.
[0006] Optionally, extracting graph structure information from the target source code to obtain a composite graph includes: Performing static analysis on the target source code to generate an abstract syntax tree and a program dependency graph; Selecting intersections from the abstract syntax tree and the statement-level nodes of the program dependency graph as nodes of the composite graph; extracting edges of a composite graph from the program dependency graph; According to the program dependency graph, a context code statement having a dependency relationship with the target node is obtained, and combined with a code statement corresponding to the target node as a first attribute; Extracting a subtree corresponding to the target node from the abstract syntax tree as a second attribute; The first attribute and the second attribute are used as node information of the target node.
[0007] Optionally, calculating the structural characteristic value of each node in the composite graph according to the edge of the composite graph includes: According to the number of edges connected to each node, the degree centrality index value of each graph node is obtained; According to the adjacency matrix of the composite graph, a Katz centrality index value of each graph node is calculated; Perform the shortest path analysis on each node to obtain the closeness centrality index value of each node; After normalizing the degree centrality index value, Katz centrality index value, and closeness centrality index value of each node, the average value is taken as the structural characteristic value of each node.
[0008] Optionally, the embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector; The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector; The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
[0009] Optionally, the detection model is a graph convolutional neural network model; the graph convolutional neural network model includes multiple convolutional layers, feature fusion layers, pooling layers and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function; The feature fusion layer is used to assign fusion weights to each convolution layer using a self-attention mechanism, and perform weighted summation on the output feature matrices of each convolution layer according to the fusion weights to obtain node fusion features; The pooling layer is used to perform a pooling operation on the node fusion features to obtain a graph feature vector; The fully connected layer is used to obtain a detection result according to the graph feature vector.
[0010] A second aspect of the present invention provides a software supply chain risk monitoring and early warning device, the device comprising: A source code preprocessing module, used to obtain the source code of the open source component that the software depends on, and preprocess the target code fragment to obtain the target source code; the target code fragment is any function segment in the source code of the open source component; A graph generation module, used to extract graph structure information from the target source code to obtain a composite graph; A structural feature extraction module, used to calculate the structural feature value of each node in the composite graph according to the edges of the composite graph; A node embedding module, used to vectorize the node information of the target node according to the pre-trained embedding representation model to obtain the node feature vector of the target node; the target node is any node in the composite graph; A graph embedding module, used for obtaining a graph embedding matrix and a structural eigenvector according to the node eigenvectors and structural eigenvalues of all nodes in the composite graph; A detection module, used to use the graph embedding matrix and the structural feature vector as inputs of a preset detection model to obtain a detection result; the detection result includes code vulnerabilities and code security; The early warning module is used to generate early warning information if the detection result shows that there is a vulnerability in the code.
[0011] Optionally, the graph generation module includes: A relationship extraction module, used for performing static analysis on the target source code to generate an abstract syntax tree and a program dependency graph; A node determination module, used to select an intersection from the abstract syntax tree and the statement-level nodes of the program dependency graph as a node of the composite graph; An edge determination module, used for extracting edges of a composite graph from the program dependency graph; A first attribute extraction module is used to obtain, according to the program dependency graph, a context code statement having a dependency relationship with the target node, and combine the context code statement corresponding to the target node as a first attribute; A second attribute extraction module, used to extract a subtree corresponding to the target node from the abstract syntax tree as a second attribute; The node information determination module is used to use the first attribute and the second attribute as the node information of the target node.
[0012] Optionally, the structural feature extraction module includes: The degree centrality calculation module is used to obtain the degree centrality index value of each graph node according to the number of edges connected to each node; A Katz centrality calculation module, used to calculate the Katz centrality index value of each graph node according to the adjacency matrix of the composite graph; The closeness centrality calculation module is used to perform the shortest path analysis on each node and obtain the closeness centrality index value of each node; The structural information synthesis module is used to normalize the degree centrality index value, Katz centrality index value and closeness centrality index value of each node, and take the average value as the structural characteristic value of each node.
[0013] Optionally, the embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector; The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector; The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
[0014] Optionally, the detection model is a graph convolutional neural network model; the graph convolutional neural network model includes multiple convolutional layers, feature fusion layers, pooling layers and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function; The feature fusion layer is used to assign fusion weights to each convolution layer using a self-attention mechanism, and perform weighted summation on the output feature matrices of each convolution layer according to the fusion weights to obtain node fusion features; The pooling layer is used to perform a pooling operation on the node fusion features to obtain a graph feature vector; The fully connected layer is used to obtain a detection result according to the graph feature vector.
[0015] Beneficial effects of the present invention: The present invention proposes a software supply chain risk monitoring and early warning method, which includes: obtaining the source code of the open source component that the software depends on, and preprocessing the target code fragment to obtain the target source code; the target code fragment is any function segment in the source code of the open source component; extracting the graph structure information of the target source code to obtain a composite graph; calculating the structural feature value of each node in the composite graph according to the edge of the composite graph; vectorizing the node information of the target node according to a pre-trained embedding representation model to obtain the node feature vector of the target node; the target node is any node in the composite graph; obtaining a graph embedding matrix and a structural feature vector according to the node feature vectors and structural feature values of all nodes in the composite graph; using the graph embedding matrix and the structural feature vector as the input of a preset detection model to obtain a detection result; the detection result includes code vulnerability and code security; if the detection result is that the code vulnerability exists, then generating early warning information.
[0016] By extracting the graph structure information of the source code, the grammatical structure and semantic information of the code can be better preserved, thereby improving the comprehensiveness and accuracy of detection. By calculating the structural feature value of each node, the model can better identify key nodes and further improve the accuracy of vulnerability detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be further described below in conjunction with the accompanying drawings.
[0018] Figure 1 A flowchart of a software supply chain risk monitoring and early warning method is provided for an embodiment of the present invention; Figure 2A structural diagram of a software supply chain risk monitoring and early warning device is provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The embodiment of the present invention provides a software supply chain risk monitoring and early warning method. Figure 1 , Figure 1 A flowchart of a software supply chain risk monitoring and early warning method provided by an embodiment of the present invention. The method comprises the following steps: S101, obtaining the source code of the open source components that the software depends on, and preprocessing the target code fragments to obtain the target source code.
[0021] S102, extracting graph structure information from the target source code to obtain a composite graph.
[0022] S103, calculating the structural characteristic value of each node in the composite graph according to the edges of the composite graph.
[0023] S104, according to the pre-trained embedding representation model, the node information of the target node is vectorized to obtain a node feature vector of the target node.
[0024] S105, obtaining a graph embedding matrix and a structural eigenvector according to the node eigenvectors and structural eigenvalues of all nodes in the composite graph.
[0025] S106, using the graph embedding matrix and the structural feature vector as input of a preset detection model to obtain a detection result.
[0026] S107, if the detection result shows that there is a vulnerability in the code, a warning message is generated.
[0027] Among them, the target code fragment is any function segment in the source code of the open source component; the target node is any node in the composite graph; the detection results include code vulnerabilities and code security.
[0028] Based on a software supply chain risk monitoring and early warning method provided by an embodiment of the present invention, by extracting graph structure information from source code, the grammatical structure and semantic information of the code can be better preserved, thereby improving the comprehensiveness and accuracy of detection. By calculating the structural feature value of each node, it helps the model to better identify key nodes and further improve the accuracy of vulnerability detection.
[0029] In one implementation, preprocessing may include: performing standardized character replacement on the function names and variable names of the target code snippet; deleting irrelevant characters in the target code snippet; irrelevant characters include comments, blank lines, redundant spaces, and line breaks. Performing standardized character replacement on function names and variable names helps avoid the impact caused by differences in naming styles among different developers. Different projects may have different naming styles. After standardization, the model is more easily generalized to different code bases, thereby improving the adaptability and stability of vulnerability detection.
[0030] In one embodiment, step S102 includes: Step 1: Perform static analysis on the target source code to generate an abstract syntax tree and program dependency graph.
[0031] Step 2: Select the intersection from the statement-level nodes of the abstract syntax tree and the program dependency graph as the nodes of the composite graph.
[0032] Step 3: Extract the edges of the composite graph from the program dependency graph.
[0033] Step 4: According to the program dependency graph, a context code statement having a dependency relationship with the target node is obtained, and combined with the code statement corresponding to the target node, it is used as the first attribute.
[0034] Step 5: Extract the subtree corresponding to the target node from the abstract syntax tree as the second attribute.
[0035] Step six: Use the first attribute and the second attribute as node information of the target node.
[0036] In one implementation, the abstract syntax tree (AST) captures the grammatical structure of the code, and the program dependency graph (PDG) reveals the dependencies of the code. The construction of this composite graph can more comprehensively reflect the logical relationships and behaviors of the code.
[0037] In one implementation, each node in the composite graph contains not only its own information, but also incorporates the code context information associated with it. This approach helps capture code dependencies across statements and improves the effectiveness of vulnerability detection, especially when there are critical vulnerabilities in the data flow or control flow. The subtree extracted from the AST as the second attribute can retain the local grammatical structure information of the target node, which helps to identify specific code patterns or potential vulnerability scenarios. By combining the context code statement (first attribute) and the grammar subtree (second attribute) as the node information of the target node, a multi-dimensional node representation can be constructed. Compared with simple grammatical or dependency information, this multi-dimensional node information improves the richness of the code representation and can provide more information support for subsequent feature extraction and detection models.
[0038] In one embodiment, step S103 includes: Step 1: According to the number of edges connected to each node, the degree centrality index value of each graph node is obtained.
[0039] Step 2: Calculate the Katz centrality index value of each graph node based on the adjacency matrix of the composite graph.
[0040] Step three, perform shortest path analysis on each node to obtain the closeness centrality index value of each node.
[0041] Step 4: Normalize the degree centrality index value, Katz centrality index value, and closeness centrality index value of each node, and take the average value as the structural characteristic value of each node.
[0042] In one implementation, degree centrality emphasizes the direct connection of local adjacent nodes, while Katz centrality and closeness centrality capture the global influence and global reachability of nodes, respectively. This multi-dimensional centrality analysis method can help more accurately identify key nodes in the code, especially those that play an important role in dependencies and code execution paths.
[0043] In one implementation, during the calculation of the three centrality index values, the directionality of the edges of the composite graph is ignored and the calculation is performed as an undirected graph. The degree centrality index value is: , in, is the degree centrality index value of node i; is node i; is the degree of node i.
[0044] The Katz centrality index value is: , in, is the Katz centrality index value of node i; is the value of the element in the i-th row and j-th column of the undirected graph adjacency matrix; N is the number of nodes; is the attenuation coefficient, which is less than the inverse of the maximum eigenvalue of the adjacency matrix and greater than 0; is the bias constant, which can be 0.5. It can be calculated iteratively, for example, the Katz centrality index values of the initialized nodes are , and then perform iterative operations until convergence.
[0045] The closeness centrality index value is: , in, is the closeness centrality index value of node i; is the shortest distance from node i to node j.
[0046] In one embodiment, the embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector.
[0047] The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector.
[0048] The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
[0049] In one implementation, the first embedding module and the second embedding module can have embedded models such as codeBERT model or GloVe. By extracting information of different structures through a preset data set, a corpus for training the embedding model is obtained. This realizes the conversion of text sequence data and tree structure data into vector form.
[0050] In one embodiment, the step detection model is a graph convolutional neural network model; the graph convolutional neural network model includes multiple convolutional layers, feature fusion layers, pooling layers and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function.
[0051] The feature fusion layer is used to assign fusion weights to each convolution layer using the self-attention mechanism, and perform weighted summation on the output feature matrices of each convolution layer according to the fusion weights to obtain the node fusion features.
[0052] The pooling layer is used to perform pooling operations on the node fusion features to obtain the graph feature vector.
[0053] The fully connected layer is used to obtain the detection results based on the graph feature vector.
[0054] In one implementation, multi-layer convolutional layers can gradually aggregate information from neighboring nodes by updating node features layer by layer, thereby deeply mining dependencies and potential vulnerabilities in the code. In multi-layer convolutional layers, by introducing the structural feature matrix obtained by expanding the structural feature vector, the high-dimensional features of the nodes are updated, avoiding the problem of over-smoothing of key node features.
[0055] In one implementation, the size of the structural feature matrix is kept consistent with the size of the output feature matrix of each layer during the calculation process. For example, if the structural feature vector is a column vector of dimension N (a matrix of size N×1), and the output feature matrix of the kth layer is of size N×16, then The size of is N×16, and each column has the same value as the structural feature vector.
[0056] In one implementation, The calculation method is: , where A is the adjacency matrix of the composite graph; D is the out-degree matrix.
[0057] In one implementation, multiple convolutional layers extract and fuse node features layer by layer, use the self-attention mechanism to assign weights to each convolutional layer, and fuse features from different convolutional layers through weighted summation. The model can comprehensively analyze local and global features in the code, thereby improving the accuracy of vulnerability detection.
[0058] In one implementation, the activation function may be ReLU.
[0059] The embodiment of the present invention provides a software supply chain risk monitoring and early warning method. Figure 2 , Figure 2 This is a structural diagram of a software supply chain risk monitoring and early warning device provided by an embodiment of the present invention. The device includes: The source code preprocessing module is used to obtain the source code of the open source components that the software depends on, and preprocess the target code fragments to obtain the target source code.
[0060] The graph generation module is used to extract graph structure information from the target source code to obtain a composite graph.
[0061] The structural feature extraction module is used to calculate the structural feature value of each node in the composite graph according to the edges of the composite graph.
[0062] The node embedding module is used to vectorize the node information of the target node according to the pre-trained embedding representation model to obtain the node feature vector of the target node.
[0063] The graph embedding module is used to obtain the graph embedding matrix and structural feature vector according to the node feature vectors and structural feature values of all nodes in the composite graph.
[0064] The detection module is used to use the graph embedding matrix and the structural feature vector as the input of the preset detection model to obtain the detection result.
[0065] The early warning module is used to generate early warning information if the detection result shows that there is a vulnerability in the code.
[0066] Among them, the target code fragment is any function segment in the source code of the open source component; the target node is any node in the composite graph; the detection results include code vulnerabilities and code security.
[0067] Based on the software supply chain risk monitoring and early warning device provided by the embodiment of the present invention, by extracting the graph structure information of the source code, the grammatical structure and semantic information of the code can be better retained, thereby improving the comprehensiveness and accuracy of the detection. By calculating the structural feature value of each node, it helps the model to better identify key nodes and further improve the accuracy of vulnerability detection.
[0068] In one embodiment, the graph generation module includes: The relation extraction module is used to perform static analysis on the target source code and generate an abstract syntax tree and program dependency graph.
[0069] The node determination module is used to select the intersection from the statement-level nodes of the abstract syntax tree and the program dependency graph as the nodes of the composite graph.
[0070] The edge determination module is used to extract the edges of the composite graph from the program dependency graph.
[0071] The first attribute extraction module is used to obtain the context code statement having a dependency relationship with the target node according to the program dependency graph, and combine it with the code statement corresponding to the target node as the first attribute.
[0072] The second attribute extraction module is used to extract the subtree corresponding to the target node from the abstract syntax tree as the second attribute.
[0073] The node information determination module is used to use the first attribute and the second attribute as the node information of the target node.
[0074] In one embodiment, the structural feature extraction module includes: The degree centrality calculation module is used to obtain the degree centrality index value of each graph node according to the number of edges connected to each node.
[0075] The Katz centrality calculation module is used to calculate the Katz centrality index value of each graph node according to the adjacency matrix of the composite graph.
[0076] The closeness centrality calculation module is used to perform the shortest path analysis on each node and obtain the closeness centrality index value of each node.
[0077] The structural information synthesis module is used to normalize the degree centrality index value, Katz centrality index value and closeness centrality index value of each node, and take the average value as the structural characteristic value of each node.
[0078] In one embodiment, the embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector.
[0079] The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector.
[0080] The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
[0081] In one embodiment, the graph convolutional neural network model includes multiple layers of convolutional layers, feature fusion layers, pooling layers, and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function; The feature fusion layer is used to assign fusion weights to each convolution layer using the self-attention mechanism, and perform weighted summation of the output feature matrices of each convolution layer according to the fusion weights to obtain the node fusion features; The pooling layer is used to perform pooling operations on the node fusion features to obtain the graph feature vector; The fully connected layer is used to obtain the detection results based on the graph feature vector.
[0082] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A software supply chain risk monitoring and early warning method, characterized in that: The method comprises: Obtaining the source code of the open source component that the software depends on, and preprocessing the target code fragment to obtain the target source code; the target code fragment is any function segment in the source code of the open source component; Extracting graph structure information from the target source code to obtain a composite graph; According to the edges of the composite graph, a structural characteristic value of each node in the composite graph is calculated; According to the pre-trained embedding representation model, the node information of the target node is vectorized to obtain a node feature vector of the target node; the target node is any node in the composite graph; Obtaining a graph embedding matrix and a structural eigenvector according to the node eigenvectors and structural eigenvalues of all nodes in the composite graph; The graph embedding matrix and the structural feature vector are used as inputs of a preset detection model to obtain a detection result; the detection result includes code vulnerabilities and code security; If the detection result shows that there is a vulnerability in the code, an early warning message is generated.
2. A software supply chain risk monitoring and early warning method according to claim 1, characterized in that: The extracting graph structure information from the target source code to obtain a composite graph includes: Performing static analysis on the target source code to generate an abstract syntax tree and a program dependency graph; Selecting intersections from the abstract syntax tree and the statement-level nodes of the program dependency graph as nodes of the composite graph; extracting edges of a composite graph from the program dependency graph; According to the program dependency graph, a context code statement having a dependency relationship with the target node is obtained, and combined with a code statement corresponding to the target node as a first attribute; Extracting a subtree corresponding to the target node from the abstract syntax tree as a second attribute; The first attribute and the second attribute are used as node information of the target node.
3. A software supply chain risk monitoring and early warning method according to claim 1, characterized in that: The step of calculating the structural characteristic value of each node in the composite graph according to the edge of the composite graph comprises: According to the number of edges connected to each node, the degree centrality index value of each graph node is obtained; According to the adjacency matrix of the composite graph, a Katz centrality index value of each graph node is calculated; Perform the shortest path analysis on each node to obtain the closeness centrality index value of each node; After normalizing the degree centrality index value, Katz centrality index value, and closeness centrality index value of each node, the average value is taken as the structural characteristic value of each node.
4. A software supply chain risk monitoring and early warning method according to claim 2, characterized in that: The embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector; The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector; The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
5. A software supply chain risk monitoring and early warning method according to claim 1, characterized in that: The detection model is a graph convolutional neural network model; the graph convolutional neural network model includes multiple convolutional layers, feature fusion layers, pooling layers and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function; The feature fusion layer is used to assign fusion weights to each convolution layer using a self-attention mechanism, and perform weighted summation on the output feature matrices of each convolution layer according to the fusion weights to obtain node fusion features; The pooling layer is used to perform a pooling operation on the node fusion features to obtain a graph feature vector; The fully connected layer is used to obtain a detection result according to the graph feature vector.
6. A software supply chain risk monitoring and early warning device, characterized in that: The device comprises: A source code preprocessing module, used to obtain the source code of the open source component that the software depends on, and preprocess the target code fragment to obtain the target source code; the target code fragment is any function segment in the source code of the open source component; A graph generation module, used to extract graph structure information from the target source code to obtain a composite graph; A structural feature extraction module, used to calculate the structural feature value of each node in the composite graph according to the edges of the composite graph; A node embedding module, used to vectorize the node information of the target node according to the pre-trained embedding representation model to obtain the node feature vector of the target node; the target node is any node in the composite graph; A graph embedding module, used for obtaining a graph embedding matrix and a structural eigenvector according to the node eigenvectors and structural eigenvalues of all nodes in the composite graph; A detection module, used to use the graph embedding matrix and the structural feature vector as inputs of a preset detection model to obtain a detection result; the detection result includes code vulnerabilities and code security; The early warning module is used to generate early warning information if the detection result shows that there is a vulnerability in the code.
7. A software supply chain risk monitoring and early warning device according to claim 6, characterized in that: The graph generation module comprises: A relationship extraction module, used for performing static analysis on the target source code to generate an abstract syntax tree and a program dependency graph; A node determination module, used to select an intersection from the abstract syntax tree and the statement-level nodes of the program dependency graph as a node of the composite graph; An edge determination module, used for extracting edges of a composite graph from the program dependency graph; A first attribute extraction module is used to obtain, according to the program dependency graph, a context code statement having a dependency relationship with the target node, and combine the context code statement corresponding to the target node as a first attribute; A second attribute extraction module, used to extract a subtree corresponding to the target node from the abstract syntax tree as a second attribute; The node information determination module is used to use the first attribute and the second attribute as the node information of the target node.
8. A software supply chain risk monitoring and early warning device according to claim 6, characterized in that: The structural feature extraction module comprises: The degree centrality calculation module is used to obtain the degree centrality index value of each graph node according to the number of edges connected to each node; A Katz centrality calculation module, used to calculate the Katz centrality index value of each graph node according to the adjacency matrix of the composite graph; The closeness centrality calculation module is used to perform the shortest path analysis on each node and obtain the closeness centrality index value of each node; The structural information synthesis module is used to normalize the degree centrality index value, Katz centrality index value and closeness centrality index value of each node, and take the average value as the structural characteristic value of each node.
9. A software supply chain risk monitoring and early warning device according to claim 7, characterized in that: The embedding representation model includes a first embedding module, a second embedding module and an attribute feature fusion module; wherein: The first embedding module is used to vectorize the first attribute of the target node to obtain a first feature vector; The second embedding module is used to vectorize the second attribute of the target node to obtain a second feature vector; The attribute feature fusion module is used to concatenate the first feature vector and the second feature vector to obtain a node feature vector.
10. A software supply chain risk monitoring and early warning device according to claim 6, characterized in that: The detection model is a graph convolutional neural network model; the graph convolutional neural network model includes multiple convolutional layers, feature fusion layers, pooling layers and fully connected layers; wherein: Multi-layer convolutional layers are used to update node features: , in, is the output feature matrix of the k+1th layer; is the output feature matrix of the kth layer; It is the structural feature matrix obtained by expanding the structural feature vector; is the normalized adjacency matrix; is the weight matrix of the k+1th layer; Representation Matrix and matrix Perform element-wise multiplication; is the activation function; The feature fusion layer is used to assign fusion weights to each convolution layer using a self-attention mechanism, and perform weighted summation on the output feature matrices of each convolution layer according to the fusion weights to obtain node fusion features; The pooling layer is used to perform a pooling operation on the node fusion features to obtain a graph feature vector; The fully connected layer is used to obtain a detection result according to the graph feature vector.
Citation Information
Patent Citations
Vulnerability detection method based on code data stream enhanced large model
CN118246029A
Graph neural network simplified interpretation method based on proxy model
CN118940815A
Source code vulnerability detection method and system based on adaptive graph neural network
CN119272275A
Source code vulnerability detection method and device based on fusion code attribute graph and electronic equipment
CN119646817A
Multi-lingual line-of-code completion system
US20210034335A1