Software material risk assessment method based on hypergraph learning
Through the software material risk assessment method based on hypergraph learning, the problem of difficult identification of software material safety risks in the software supply chain is solved, high-precision risk identification and evaluation is achieved, and system security is improved.
Patent Information
- Application Number
- CN202510084410.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art ignores the security risks carried and transmitted by the software bill of materials (SBOM) during the software life cycle, making it difficult to identify and evaluate risks in the software supply chain.
A software material risk assessment method based on hypergraph learning is proposed. By constructing an HGL-SSCR model, a preprocessor is used to obtain network entity data, model simple graphs, extract fragile correlations to build hypergraphs, aggregate nodes and hyperedge features, update the centroid vertex features, and finally output the fragility classification label of network entities.
It improves the accuracy of identifying software supply chain risks, can effectively identify vulnerability risks hidden in software materials, improves system security, and demonstrates excellent macro accuracy, macro recall and macro F1 index in practical applications.
Smart Images

Figure CN120017332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network entity vulnerability risk assessment, and in particular to a software material risk assessment method. Background Art
[0002] Network entity vulnerability is a weak link in an asset or asset group that may be exploited by threats to cause damage. The goal of the network entity vulnerability risk assessment task is to obtain the potential vulnerability risks in network entities from the active defense process. As an important entity of network services, the construction of software systems increasingly relies on the integration of a large number of third-party software components. Although this trend has promoted technological innovation, it has also greatly increased the risk of attacks on the software supply chain. Because it is impossible to obtain a complete and real-time list of software components, clone vulnerabilities in the software supply chain are hidden and difficult to identify. Services based on network entity vulnerability risk assessment are widely used in network security incident monitoring and response, network performance optimization, IoT device security assessment, and enterprise security management.
[0003] Traditional network entity direct vulnerability assessment algorithms are divided into two types: matching based on vulnerability databases and analysis of key network entity vulnerabilities. Based on vulnerability database matching, the vulnerability database maintains a set of mappings between vulnerabilities and service configurations, including service features such as operating systems, version numbers, and port numbers, and provides vulnerability query services. Although the above two methods have achieved certain results in the field of network entity vulnerability assessment, they also face challenges such as 0day vulnerability processing, large-scale repair costs, and security risks introduced by third-party components.
[0004] Indirect vulnerability assessment algorithm for network entities based on software supply chain. In recent years, the risks of software supply chain have increased in cybersecurity events such as attacks and frauds. Vulnerabilities from transitive dependencies are more difficult to predict in the context of incomplete software component lists. In the face of this challenge, Andre et al. proposed a new system for assessing software supply chain risks, which overcomes the difficulty of less label data that can be learned in the software supply chain. By combining Bayesian belief network & AHP & Noisy-OR technology to identify software supply chain risks, the query efficiency in the decision-making process is effectively improved. Bilal et al. innovatively proposed a social technology framework in the field of software supply chain, which can identify social and technical risks of software supply chain by modeling software supply chain system. Robert et al. conducted fuzz testing from the practical perspective of software quality, screened vulnerabilities by using malformed input, and then designed and coded through risk assessment technology, and finally achieved the result of reducing software supply chain risks. Duan et al. proposed the tool MalOSS, which discovers abnormal data flows by comparing the behavioral history of software updates, thereby identifying malicious data packets, and finally assessing the risks of software supply chain.
[0005] However, the above methods lack a reasonable hypergraph construction scheme based on software supply chain risk labels, and ignore the security risks carried and transmitted by the software bill of materials (SBOM) during the software life cycle, making it challenging to use hypergraph learning mechanisms to transmit software supply chain risk labels on nodes. Summary of the invention
[0006] In view of the technical problem that the existing technology ignores the security risks carried and transmitted by the software bill of material (SBOM) during the software life cycle, the present invention proposes a software material risk assessment method based on hypergraph learning, which improves the accuracy of risk identification while taking into account the security risks carried and transmitted by the software bill of material.
[0007] In order to achieve the above object, the technical solution of the present invention is achieved as follows:
[0008] A software material risk assessment method based on hypergraph learning comprises the following steps:
[0009] S1: Constructing the HGL-SSCR model, including the preprocessor, hypergraph building module, hypergraph learning module and output module connected in sequence;
[0010] S2: Use the preprocessor to obtain network entity data and use the network entity data to model a simple graph;
[0011] S3: Use the hypergraph building module to extract fragile correlations from simple graphs and build a hypergraph;
[0012] S4: Use the hypergraph learning module to aggregate node features in the hypergraph to obtain hyperedge features, and aggregate hyperedge features to update centroid vertex features;
[0013] S5: Taking the centroid vertex feature vector as input, the output module is used to classify and output the vulnerability classification label of the network entity.
[0014] Further, the method of obtaining network entity data by using a preprocessor is:
[0015] S2.1: Use Ipv6toolkit to detect live IP addresses on the target IP network segment and obtain a set of live IP addresses;
[0016] S2.2: Collect surviving IP addresses based on the surviving IP addresses and use a variety of network scanning and identification methods to obtain surviving network entity measurement information;
[0017] S2.3: Use the Traceroute command to measure surviving network entities and construct the IP network topology;
[0018] S2.4: Verify the surviving network entity measurement information through the network entity resource detection and analysis platform to obtain the final network entity data.
[0019] Furthermore, the method of obtaining the measurement information of surviving network entities using multiple network scanning and identification methods is as follows: obtaining network entity attribute characteristics through Shodan, Zmap, zoomeye, Nmap, and distributed remote packet capture methods; obtaining vulnerability information through the AWVS vulnerability scanning tool, matching the vulnerability information in the vulnerability information database, and recording the vulnerability source as a third-party middleware feature;
[0020] In the process of verifying the surviving network entity measurement information through the network entity resource detection and analysis platform, a vulnerability classification label is added to each surviving network entity according to the vulnerability source in the third-party middleware characteristics; the network entity data includes network entity measurement information, IP network topology data and network entity vulnerability classification labels.
[0021] Further, the method of modeling a simple graph using network entity data in step S2 is: extracting attributes according to the obtained network entity measurement information, and encoding data according to the attribute extraction result to obtain a simple graph;
[0022] The data encoding method is:
[0023] Select the IP in the network entity measurement information as the node υ in the simple graph i =IP i , select the connection relationship between two IPs in the IP network topology data as edge e ij =IP i →IP j Node υ i By transforming the matrix Get node v i Initial node features Transformation Matrix Represents node attributes, edge e ij Through the transformation matrix Get edge e ij Initial edge features Transformation Matrix Represents the edge attribute, and the node υ i , edge ij , node υ i Initial node features and edge ij Initial edge features Encode them separately to get a simple graph g = (υ, e, h υ ,h e );
[0024] where υ represents a node set and is denoted as υ=(υ 1 , 2 , ..., υ i , ...), e is represented as a set of edges and is denoted by e = (e 12 , e 34 , ..., e ij , ...), Represents a low-dimensional initial embedding of node features in a simple graph, Representing low-dimensional initial embeddings of edge features in simple graphs.
[0025] Further, the method for extracting fragile correlations and constructing a hypergraph in step S3 is:
[0026] Use Z-score standardization to normalize the initial node features Perform normalization and calculate the features of each initial node The Euclidean distance to the neighboring nodes, select the first K nearest neighboring nodes, and get the distance to the node υ i The K neighbor nodes with the most relevant vulnerability, with node υ i As the centroid vertex, connect K neighbor nodes to form a hyperedge, and assign weights to the K neighbor nodes. The closer the distance, the higher the weight. After traversing all nodes, we finally get the hypergraph H = {υ, ε}, ε = (ε 1 , ε 2 , ...ε k , ...), the specific formula is:
[0027]
[0028] in, For each node i and the set of K selected neighbor nodes; ε k Represents a constructed hyperedge.
[0029] Furthermore, the method of aggregating node features in the hypergraph to obtain hyperedge features is:
[0030] Furthermore, a multi-layer perceptron is used to learn the transformation matrix T from the node features. After multiplying the transformation matrix T with the node features, a one-dimensional convolution is used to obtain the hyperedge features. The formula is:
[0031]
[0032] Among them, MLP() represents the multi-layer perceptron calculation function, conv() is the convolution function, is the node feature, It is a super edge feature.
[0033] Furthermore, the method of aggregating hyperedge features to update centroid vertex features is:
[0034] Transform hyperedge features using weight matrix W and neighbor hyperedge features After splicing the transformed hyperedge features, the attention score is calculated using the linear transformation matrix a, and normalized using the LeakyReLU activation function to obtain the graph attention coefficient a ij , where the hyperedge feature is the center of mass vertex υ i Hyperedge ε i Hyperedge features, neighbor hyperedge features is the center of mass vertex υ i Hyperedge ε i The neighbor hyperedge ε j Hyperedge features of; Calculate neighbor hyperedge features ; Apply activation function σ to update the centroid vertex features; Finally, update all centroid vertex feature vectors.
[0035] Furthermore, the graph attention coefficient a ij The formula is:
[0036]
[0037] Among them, LeakyReLU() is the LeakyReLU activation function, softmax() is the normalized activation function, and a T is the transposed matrix of the linear transformation matrix a, || represents the concatenation operation;
[0038] The calculation of neighbor hyperedge features The weighted sum method is:
[0039]
[0040] Among them, τ v To adjust the neighbor hyperedge features The scaling factor, is the neighbor hyperedge feature in the lth layer of the hyperedge convolution module, E i represents the hyperedge ε i The set of neighbor hyperedges of ;
[0041] The method of applying the activation function σ to update the centroid vertex features is:
[0042]
[0043] in, Represents the hyperedge features at the l+1th layer of the hyperedge convolution module.
[0044] Furthermore, the calculation process of classifying and outputting the vulnerability classification labels of network entities is as follows:
[0045]
[0046] Among them, HGL-SSCR (v i ) is the vulnerability prediction transformation matrix of each node, which is used to output the vulnerability classification label of the network entity, Sigmoid() is the activation function, BN() is the batch normalization function, W nev is the weight matrix derived from the multilayer perceptron, b nev is the coefficient of variation.
[0047] Furthermore, the training loss function used by the HGL-SSCR model is:
[0048]
[0049] Among them, Loss is the loss function, n represents the number of nodes, y i The real label of the vulnerability classification of the network entity obtained in step 2.4, is the vulnerability classification prediction value of the network entity,
[0050] The beneficial effects of the present invention are:
[0051] The method of the present invention proposes a network entity vulnerability analysis method based on the software supply chain, which can obtain the vulnerability risks hidden in the software materials and thus improve the system security.
[0052] This paper proposes a learning strategy based on a hypergraph neural network, which effectively improves the accuracy and performance of the model compared to simple graph deep learning models.
[0053] The model of the present invention is superior to the existing models in terms of macro precision, macro recall and macro F1 index, and can accurately identify the software material vulnerability labels of network entities.
[0054] In addition, the model of the present invention has been successfully applied to the vulnerability platform, providing a risk assessment chain for network entities under the security service software supply. The platform has been online for seven months, providing safe, stable and fast detection and protection services for more than 2,000 industry customers. It fully proves its important role in software material risk identification and defense. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0056] Figure 1 The present invention is a flow chart of the method.
[0057] Figure 2 This is a structural diagram of the HGL-SSCR model of the present invention.
[0058] Figure 3 This is part of the network entity measurement information extracted by the method of the present invention.
[0059] Figure 4 Implementation process diagram for hypergraph construction.
[0060] Figure 5 Implementation diagram of hypergraph learning.
[0061] Figure 6 Implement a procedural graph for vertex convolution. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] like Figure 1 As shown, a software supply chain risk assessment method based on hypergraph learning includes the following steps:
[0064] S1: Construct the HGL-SSCR model, including the preprocessor, hypergraph building module, hypergraph learning module and output module connected in sequence, such as Figure 2 shown.
[0065] The preprocessor is used to obtain network entity data and use the network entity data to model a simple graph;
[0066] The hypergraph construction module is used to extract fragile correlations from simple graphs and construct hypergraphs;
[0067] The hypergraph learning module is used to aggregate node features in the hypergraph to obtain hyperedge features, and aggregate hyperedge features to update centroid vertex features; the hypergraph learning module includes a vertex convolution module and a hyperedge convolution module connected in sequence; the vertex convolution module is used to aggregate node features in the hypergraph to obtain hyperedge features, and the hyperedge convolution module is used to aggregate hyperedge features to update centroid vertex features.
[0068] S2: Use the preprocessor to obtain network entity data and use the network entity data to model a simple graph.
[0069] To ensure the integrity of network measurement data, the acquisition steps are as follows:
[0070] S2.1: Use Ipv6toolkit to detect live IP addresses on the target IP network segment and obtain a set of live IP addresses;
[0071] S2.2: Based on the surviving IP addresses, the surviving IPs are concentrated and the surviving network entity measurement information is obtained using a variety of network scanning and identification methods. The network entity attribute characteristics such as operating system information, port list and service, web metadata, web framework, application layer service protocol information, packet size sequence, delay information, routing information, and IP rate limit are obtained through Shodan, Zmap, zoomeye, Nmap, and distributed remote packet capture methods; vulnerability information is obtained through the AWVS vulnerability scanning tool, and the vulnerability information is matched in the vulnerability information database, and the vulnerability source (software or third-party plug-in) is recorded as a third-party middleware feature (i.e., SBOM). Some network entity measurement information such as Figure 3 shown.
[0072] S2.3: Use the Traceroute command to measure surviving network entities and construct the IP network topology;
[0073] S2.4: Verify the surviving network entity measurement information through the network entity resource detection and analysis platform (DayDaymap), supplement the missing information in the surviving network entity measurement information, ensure the accuracy and completeness of the data, add vulnerability classification labels to each surviving network entity according to the vulnerability source in the SBOM, and obtain the final network entity data; the network entity data includes network entity measurement information, IP network topology data and network entity vulnerability classification labels. In order to improve the stability of network measurement data, the present invention measures multiple times and controls the measurement bandwidth at 10m / s.
[0074] Use network entity data to model a simple graph: perform attribute extraction based on the obtained network entity measurement information, perform data encoding based on the attribute extraction results, and obtain a simple graph.
[0075] Attribute extraction:
[0076] The network entity attribute features are selected as the node features. In the selection of edge attributes, similar functions are often provided to similar service objects. This embodiment pays more attention to the neighbor information around each node and the correlation between multiple nodes. The edge is only used to represent the connection relationship between two nodes.
[0077] According to the attribute extraction results, data encoding is performed to obtain a simple graph:
[0078] In the context of software supply chain, network entities often have multiple IP addresses, so the IP in the network entity measurement information is selected as the node v in the simple graph. i =IP i , select the connection relationship between two IPs in the IP network topology data as edge e ij =IP i →IP j ; In the simple graph modeling task, an edge represents the routing connection relationship between two IP addresses, and the edge is composed of a node υ i Pointing to node v j , that is, e ij =υ i → j , that is (IP i →IP j ).
[0079] Nodeυ i By transforming the matrix Get node v i Initial node features Transformation Matrix Represents node attributes, edge e ij Through the transformation matrix Get edge e ij Initial edge features Transformation Matrix Represents the edge attribute, and the node υ i , edge ij , node υ i Initial node features and edge ij Initial edge features Encode them separately to get a simple graph g = (υ, e, h υ ,h e ), where υ represents the node set and is denoted as υ = (υ 1 , 2 , .., υ i , ...)=(IP 1 , IP 2 ,...,IP i , ...), e is represented as a set of edges and is denoted by e = (e 12, e 34 , ..., e ij , ...)=(IP 1 →IP 2 , IP 3 →IP 4 , ..., IP i →IP j ...), Represents a low-dimensional initial embedding of node features in a simple graph, Representing low-dimensional initial embeddings of edge features in simple graphs.
[0080] S3: Use the hypergraph building module to extract fragile correlations from simple graphs and build a hypergraph.
[0081] In order to improve the accuracy of network entity vulnerability identification, it is necessary to extract a set of high-dimensional relationships with strong vulnerability correlation to construct a hypergraph, extract vulnerability correlation and construct a hypergraph, such as Figure 4 As shown:
[0082] Use Z-score standardization to normalize the initial node features Perform normalization and calculate the features of each initial node The Euclidean distance to the neighboring nodes, select the first K nearest neighboring nodes, and get the distance to the node υ i The K neighbor nodes with the most relevant vulnerability, with node υ i As the centroid vertex, connect K neighbor nodes to form a hyperedge, and assign weights to the K neighbor nodes. The closer the distance, the higher the weight. After traversing all nodes, we finally get the hypergraph H = {υ, ε}, ε = (ε 1 , ε 2 , ...ε k , ...), the specific formula is:
[0083]
[0084] in, For each node i and the set of K selected neighbor nodes; ε k Represents a constructed hyperedge.
[0085] S4: Use the hypergraph learning module to aggregate node features in the hypergraph to obtain hyperedge features, and aggregate hyperedge features to update centroid vertex features, such as Figure 5 shown.
[0086] The node features in the hypergraph are aggregated through the vertex convolution module to obtain the hyperedge features, and the hyperedge features are aggregated through the hyperedge convolution module to update the centroid vertex features.
[0087] Vertex convolution module: such as Figure 6As shown in the figure, a multi-layer perceptron is used to learn the transformation matrix T from the node features. After multiplying the transformation matrix T with the node features, a one-dimensional convolution is used to obtain the hyperedge features, as shown in the formula:
[0088]
[0089] Among them, MLP() represents the multi-layer perceptron calculation function, conv() is the convolution function, is the node feature, It is a hyperedge feature, that is, the node features around the centroid vertex are aggregated as the hyperedge feature of the hyperedge where the centroid vertex is located.
[0090] Hyperedge convolution module: Use weight matrix W to transform hyperedge features and neighbor hyperedge features After splicing the transformed hyperedge features, the attention score is calculated using the linear transformation matrix a, and normalized using the LeakyReLU activation function to obtain the graph attention coefficient a ij , where the hyperedge feature is the center of mass vertex υ i Hyperedge ε i Hyperedge features, neighbor hyperedge features is the center of mass vertex υ i Hyperedge ε i The neighbor hyperedge ε j The hyperedge feature of is as follows:
[0091]
[0092] Among them, LeakyReLU() is the LeakyReLU activation function, softmax() is the normalized activation function, and a T is the transposed matrix of the linear transformation matrix a, and || represents the concatenation operation.
[0093] Calculate neighbor hyperedge features The weighted sum of:
[0094]
[0095] Among them, τ v To adjust the neighbor hyperedge features The scaling factor, is the neighbor hyperedge feature in the lth layer of the hyperedge convolution module, E i represents the hyperedge ε i The set of neighbor hyperedges of .
[0096] Apply the activation function σ to update the centroid vertex features:
[0097]
[0098] in, Represents the hyperedge features at the l+1th layer of the hyperedge convolution module.
[0099] Finally, all centroid vertex feature vectors are updated:
[0100] S5: Taking the centroid vertex feature vector as input, the output module is used to classify and output the vulnerability classification label of the network entity.
[0101] The calculation process of classifying and outputting the vulnerability classification label of the network entity is as follows:
[0102]
[0103] Among them, HGL-SSCR (v i ) is the vulnerability prediction transformation matrix of each node, which is used to output the vulnerability classification label of the network entity. Sigmoid is the activation function. BN() is the batch normalization function, which is used to deal with the overfitting problem. Wnev is the weight matrix derived from the multi-layer perceptron (MLP). bnev is the bias coefficient. Sigmoid() is the activation function.
[0104] The algorithm training uses the stochastic gradient descent of the Adam optimizer to reduce the loss between the predicted value and the true value. The model training loss of HGL-SSCR is:
[0105]
[0106] Among them, Loss is the loss function, n represents the number of nodes, y i The real label of the vulnerability classification of the network entity obtained in step 2.4, is the vulnerability classification prediction value of the network entity, that is
[0107] Experimental results:
[0108] The present invention divides the data set by region in a total of 17,636 server IP addresses in three regions, and randomly selects ratio = 60% of the server IP addresses in each region as the training set, ratio = 20% of the server IP addresses as the validation set, and ratio = 20% of the server IP addresses as the test set. The training set of the Shanghai experimental area in China includes 3,760 server IP addresses, the training set of the Hong Kong experimental area in China includes 4,485 server IP addresses, and the training set of each New York experimental area includes 2,336 server IP addresses. The results of the vulnerability risk classification of network entities indirectness in the three regions are shown in Table 4 below: The model of the present invention is superior to the existing models in macro precision, macro recall and macro F1 index.
[0109] Table 4
[0110]
[0111]
[0112] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A software material risk assessment method based on hypergraph learning, characterized in that: Includes steps: S1: Constructing the HGL-SSCR model, including the preprocessor, hypergraph building module, hypergraph learning module and output module connected in sequence; S2: Use the preprocessor to obtain network entity data and use the network entity data to model a simple graph; S3: Use the hypergraph building module to extract fragile correlations from simple graphs and build a hypergraph; S4: Use the hypergraph learning module to aggregate node features in the hypergraph to obtain hyperedge features, and aggregate hyperedge features to update centroid vertex features; S5: Taking the centroid vertex feature vector as input, the output module is used to classify and output the vulnerability classification label of the network entity.
2. The software material risk assessment method based on hypergraph learning according to claim 1 is characterized in that: The method of obtaining network entity data by using a preprocessor is: S2.1: Use Ipv6toolkit to detect live IP addresses on the target IP network segment and obtain a set of live IP addresses; S2.2: Collect surviving IP addresses based on the surviving IP addresses and use a variety of network scanning and identification methods to obtain surviving network entity measurement information; S2.3: Use the Traceroute command to measure surviving network entities and construct the IP network topology; S2.4: Verify the surviving network entity measurement information through the network entity resource detection and analysis platform to obtain the final network entity data.
3. The software material risk assessment method based on hypergraph learning according to claim 2 is characterized in that: The method of using multiple network scanning and identification methods to obtain measurement information of surviving network entities is as follows: obtaining network entity attribute characteristics through Shodan, Zmap, zoomeye, Nmap and distributed remote packet capture methods; obtaining vulnerability information through AWVS vulnerability scanning tools, matching the vulnerability information in the vulnerability information database, and recording the vulnerability source as a third-party middleware feature; In the process of verifying the surviving network entity measurement information through the network entity resource detection and analysis platform, a vulnerability classification label is added to each surviving network entity according to the vulnerability source in the third-party middleware characteristics; the network entity data includes network entity measurement information, IP network topology data and network entity vulnerability classification labels.
4. The software material risk assessment method based on hypergraph learning according to claim 2 or 3 is characterized in that: The method of modeling a simple graph using network entity data in step S2 is: extracting attributes based on the obtained network entity measurement information, and encoding data based on the attribute extraction result to obtain a simple graph; The data encoding method is: Select the IP in the network entity measurement information as the node υ in the simple graph i =IP i , select the connection relationship between two IPs in the IP network topology data as edge e ij =IP i →IP j Node υ i By transforming the matrix Get node v i Initial node features Transformation Matrix Represents node attributes, edge e ij Through the transformation matrix Get edge e ij Initial edge features Transformation Matrix Represents the edge attribute, and the node υ i , edge ij , node υ i Initial node features and edge ij Initial edge features Encode them separately to get a simple graph g = (υ, e, h u ,h e ); where υ represents a node set and is denoted as υ=(υ1,υ2,...,υ i , ...), e is represented as a set of edges and is denoted by e = (e 12 , e 34 , ..., e ij , ...), Represents a low-dimensional initial embedding of node features in a simple graph, Representing low-dimensional initial embeddings of edge features in simple graphs.
5. The software material risk assessment method based on hypergraph learning according to claim 4 is characterized in that: The method for extracting fragile correlations and constructing a hypergraph in step S3 is: Use Z-score standardization to normalize the initial node features Perform normalization and calculate the features of each initial node The Euclidean distance to the neighboring nodes, select the first K nearest neighboring nodes, and get the distance to node v i The K neighbor nodes with the most relevant vulnerability, with node v i As the centroid vertex, connect K neighbor nodes to form a hyperedge, and assign weights to the K neighbor nodes. The closer the distance, the higher the weight. After traversing all nodes, we finally get the hypergraph H = {v, ε}, ε = (ε1, ε2, ... εk , ...), the specific formula is: in, For each node i and the set of K selected neighbor nodes; ε k Represents a constructed hyperedge.
6. The software material risk assessment method based on hypergraph learning according to claim 5 is characterized in that: The method of aggregating node features in a hypergraph to obtain hyperedge features is: The transformation matrix T is learned from the node features using a multi-layer perceptron. After multiplying the transformation matrix T with the node features, the hyperedge features are obtained using a one-dimensional convolution. The formula is: Among them, MLP() represents the multi-layer perceptron calculation function, conv() is the convolution function, is the node feature, It is a super edge feature.
7. The software material risk assessment method based on hypergraph learning according to claim 6 is characterized in that: The method of aggregating hyperedge features to update centroid vertex features is: Transform hyperedge features using weight matrix W and neighbor hyperedge features After splicing the transformed hyperedge features, the attention score is calculated using the linear transformation matrix a, and normalized using the LeakyReLU activation function to obtain the graph attention coefficient a ij , where the hyperedge feature is the center of mass vertex υ i Hyperedge ε i Hyperedge features, neighbor hyperedge features is the center of mass vertex υ i Hyperedge ε i The neighbor hyperedge ε j The hyperedge features of Calculate neighbor hyperedge features The weighted sum of ; Apply activation function σ to update the centroid vertex features; Finally, all centroid vertex feature vectors are updated.
8. The software material risk assessment method based on hypergraph learning according to claim 7 is characterized in that: The graph attention coefficient a ij The formula is: Among them, LeakyReLU() is the LeakyReLU activation function, softmax() is the normalized activation function, and a T is the transposed matrix of the linear transformation matrix a, || represents the concatenation operation; The calculation of neighbor hyperedge features The weighted sum method is: Among them, τ v To adjust the neighbor hyperedge features The scaling factor, is the neighbor hyperedge feature in the lth layer of the hyperedge convolution module, E i represents the hyperedge ε i The set of neighbor hyperedges of ; The method of applying the activation function σ to update the centroid vertex features is: in, Represents the hyperedge features at the l+1th layer of the hyperedge convolution module.
9. The software material risk assessment method based on hypergraph learning according to claim 7 is characterized in that: The calculation process of classifying and outputting the vulnerability classification labels of network entities is as follows: Among them, HGL-SSCR (v i ) is the vulnerability prediction transformation matrix of each node, which is used to output the vulnerability classification label of the network entity, Sigmoid() is the activation function, BN() is the batch normalization function, W nev is the weight matrix derived from the multilayer perceptron, b nev is the coefficient of variation.
10. The software material risk assessment method based on hypergraph learning according to claim 9 is characterized in that: The training loss function used by the HGL-SSCR model is: Among them, Loss is the loss function, n represents the number of nodes, y i The real label of the vulnerability classification of the network entity obtained in step 2.4, is the vulnerability classification prediction value of the network entity,