Key node mining method and system for engineering project complex network
By using graph-based neural network methods in complex engineering projects, the project network diagram and perturbation diagram are constructed, and the node loss values are calculated are sorted with importance, which solves the problem of insufficient identification accuracy of key nodes in the existing technology, and achieves higher identification accuracy and model robustness.
Patent Information
- Application Number
- CN202510145656.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively utilize global topology information and node attribute information in complex engineering networks, resulting in insufficient accuracy of key node identification.
Using a graph neural network-based method, by constructing project network diagrams and perturbation diagrams, combining the encoder and decoder of graph neural network for embedding learning and reconstruction, the loss values of the calculation nodes are sorted by importance.
It improves the accuracy and reliability of the identification of key nodes in complex networks, comprehensively considers the local topology structure and global network impact of nodes, and enhances the robustness and generalization capabilities of the model.
Smart Images

Figure CN120068928A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of complex network analysis for engineering projects, and particularly relates to a method and system for mining key nodes for an engineering project complex network. Background Art
[0002] As the core link in the field, the construction management of engineering projects directly affects the quality and efficiency of engineering projects, and also affects the optimization of resource allocation and the smooth completion of engineering tasks. With the continuous expansion of the scale of engineering projects, the relevance in the project network becomes increasingly complex. How to identify key nodes in the complex network to improve resource utilization efficiency and capacity support has become a key issue in the construction management of engineering projects in this field.
[0003] The key nodes in a complex network refer to the nodes that occupy a core position in the network and have an important impact on the function, efficiency, and stability of the network. For example, the key task nodes in an engineering project may play a decisive role in the progress of the entire project; the failure of core resource nodes may lead to the interruption of resource allocation; and the key capacity nodes may directly affect the satisfaction of project task requirements. Therefore, identifying and optimizing key nodes is of great significance for improving the efficiency of project implementation paths, enhancing the overall stability of the network, and improving the rationality of resource allocation.
[0004] Current key node mining algorithms are mostly based on single network topology characteristics, such as degree centrality based on local characteristics or betweenness centrality and closeness centrality based on global characteristics. Algorithms based on local characteristics usually only focus on the local information within the node neighborhood and cannot accurately reflect the importance of nodes in the overall network; while algorithms based on global characteristics, although considering the overall topology structure of the network, often ignore the direct impact of node attributes on network functions. In addition, these algorithms usually cannot combine the characteristics of complex networks and are difficult to accurately depict the complex associations between projects in engineering projects.
[0005] With the rapid development of the Graph Neural Network (GNN), the application of graph learning-based methods in complex network analysis has gradually received attention. GNN can capture both local and global characteristics of the network by aggregating node neighborhood information and, combined with node attributes, provides a new technical path for mining key nodes. However, existing GNN-based methods are mostly used for the analysis of single-layer networks and still have application limitations in the identification of key nodes in complex networks. Summary of the Invention
[0006] The object of the present invention is to propose a method and system for mining key nodes of an engineering project complex network based on a graph neural network, so as to solve the problem of insufficient utilization of global topology information and node attribute information in the importance ranking of project nodes in the prior art.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A method for mining key nodes for a complex network of engineering projects, comprising the following steps:
[0009] 1) Construct a project network diagram based on the engineering project network as the original graph, where the nodes in the graph represent project entities, the edges represent the dependency relationships between projects, and define a node feature matrix;
[0010] 2) Generate a perturbed graph by removing all the edges of the target node in the original graph and retaining the node feature matrix;
[0011] 3) Perform data sampling on the original graph to obtain positive and negative sample sets and construct a training data set;
[0012] 4) Reconstruct the perturbed graph using the encoder and decoder of the graph neural network, where the node embedding vectors are generated by the encoder, the adjacency matrix is reconstructed by the decoder, and the reconstruction ability of the decoder is trained using the above training data set;
[0013] 5) Compare the differences between the original graph and the reconstructed graph, and calculate the loss value of the nodes;
[0014] 6) Sort the nodes according to the loss values of the nodes to determine the key nodes.
[0015] Further, the steps of constructing a project network diagram according to the engineering project network in step 1) include:
[0016] Represent the engineering project network as an undirected unweighted graph, and define a node set, an edge set, and a node feature matrix;
[0017] Extract node features according to the attribute information of the nodes, and construct a node feature matrix;
[0018] Generate an adjacency matrix according to the dependency relationships between the nodes;
[0019] Normalize the node feature matrix and the adjacency matrix;
[0020] Combine the node feature matrix and the adjacency matrix to generate project network diagram data.
[0021] Further, the node features include resource consumption, project contribution degree, and project cost-effectiveness ratio.
[0022] Further, the normalization processing of the node feature matrix includes normalizing the node features, and the normalization processing of the adjacency matrix includes calculating the symmetric normalized adjacency matrix.
[0023] Furthermore, the steps of generating the perturbed graph in step 2) further include:
[0024] After removing all the edges associated with the target node, update the adjacency matrix to generate a perturbed adjacency matrix;
[0025] Combine the perturbed adjacency matrix with the node feature matrix to generate perturbed graph data.
[0026] Furthermore, the steps of data sampling for the original graph in step 3) include:
[0027] Extract the node set and edge set from the original graph, and construct a positive sample set based on the actually existing edges;
[0028] Select the edges that do not exist in the adjacency matrix as candidate negative samples to construct a candidate negative sample set;
[0029] Randomly sample the non-existent edges from the candidate negative sample set to form a negative sample set, and the size of the negative sample set is equal to that of the positive sample set;
[0030] Merge the positive sample set and the negative sample set, and assign labels to each edge to generate training data.
[0031] Furthermore, in step 4), the encoder dynamically selects GCN, GAT or GraphSAGE according to the characteristics of the engineering project network and the task requirements.
[0032] Furthermore, the steps of reconstructing the adjacency matrix through the decoder in step 4) include:
[0033] Through two multi-layer perceptrons, first perform feature mapping on the input node embeddings to compress the high-dimensional embeddings to an intermediate dimension; then map the intermediate dimension to the final output space to generate a scalar value, which represents the connection probability between nodes, predict the edge connection relationship between nodes, and generate a reconstructed adjacency matrix.
[0034] Furthermore, the node loss value calculated in step 5) is specifically the mean cross-entropy loss of each node in multiple tests, which is composed of the sum of the mean positive sample loss and the mean negative sample loss.
[0035] A key node mining system for engineering project complex networks includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0036] The beneficial effects obtained by the present invention are as follows:
[0037] 1. Improve the accuracy of identifying key nodes in complex networks: The present invention removes the edges of target nodes to generate a perturbed network graph, combines graph neural networks for embedding learning and reconstruction, and constructs a perturbation-aware graph neural network. By comparing the differences between the reconstructed adjacency matrix and the original adjacency matrix, the contribution of nodes to the network structure and function is quantified. On the basis of retaining the network topology structure and node attribute information, by perturbing the network structure and calculating the reconstruction loss of the network before and after perturbation, the degree of influence of nodes on the network function is measured, so as to realize the importance ranking of key nodes. This loss evaluation method based on graph structure changes comprehensively considers the local topology structure of nodes and the global network influence, overcomes the limitations of traditional single-topology measurement methods such as degree centrality and betweenness centrality, and improves the accuracy and reliability of key node identification.
[0038] 2. Comprehensively analyze global topology and node features: The present invention combines the node features extracted by graph neural networks with network topology features, introduces a fusion representation of node attributes and graph structures, and adopts a self-supervised learning graph perturbation mechanism to comprehensively characterize the role of nodes in complex networks. The encoder module captures node embedding features at different scales through multi-layer feature aggregation, and the decoder module further reflects the connectivity differences between nodes by reconstructing the adjacency matrix. This technology can more comprehensively identify key nodes in complex networks and has important value especially in application scenarios such as resource management and network stability analysis.
[0039] 3. Enhance the robustness and generalization ability of the model: The present invention introduces a negative sampling strategy for balancing positive and negative samples in the negative sample generation process. By randomly generating negative samples from the edges that do not exist in the network and participating in model training together with the actually existing edges (positive samples), while keeping the number of positive and negative samples balanced, it prevents training bias caused by data imbalance. This negative sampling mechanism enables the model to accurately learn the importance information of nodes under diverse network structures, enhancing the robustness and generalization ability of the model in complex network environments.
[0040] 4. Wide range of application scenarios: The present invention can be widely applied to multiple scenarios such as engineering project task scheduling, resource optimization allocation, and network stability analysis, providing efficient and reliable technical support for the construction management of engineering projects in the field. Brief Description of the Drawings
[0041] Figure 1 is a brief flowchart of the key node mining method for complex networks of engineering projects in the embodiment.
[0042] Figure 2 is a specific flowchart of the key node mining method for complex networks of engineering projects in the embodiment.
[0043] Figure 3 is a schematic diagram of a project network graph. Detailed implementation manners
[0044] To make the technical features, advantages or technical effects in the above technical solutions of the present invention more obvious and understandable, the following will be described in detail with reference to embodiments.
[0045] The embodiment of the present invention specifically provides a method for mining key nodes for a complex network of engineering projects. As shown in the process of Figure 1 and Figure 2 , by designing a graph neural network with perturbation perception, the structural information and node features of the complex network can be effectively integrated, and the accuracy and reliability of key node recognition can be improved. The specific implementation steps of this method are as follows:
[0046] Step 1: Construct a project network graph
[0047] Represent the engineering project network as a project network graph, which is an undirected and unweighted graph. The nodes in the graph (also known as project nodes) represent project entities, and the edges represent the dependency relationships between projects. By analyzing the node attribute information and topological structure, a node feature matrix and an adjacency matrix are constructed as the input of the algorithm. The specific processing process is as follows:
[0048] 1) Represent the engineering project network: Represent the engineering project network as an undirected and unweighted graph g=(A, X)=(V, E, X). Wherein, V=(v 1 , v 2 ,... v N ) is the set of nodes, representing project entities; E={e ij |v i , v j ∈V} is the set of edges, representing the dependency relationships between nodes. If there is a dependency between two nodes v i and v j , there is an edge e ij ∈E. X represents the node feature matrix, which is composed of multi-dimensional attribute features.
[0049] 2) Node feature extraction: Analyze the inherent attribute information of node v i , and extract multi-dimensional features, including resource consumption, project contribution, project cost-effectiveness ratio, etc., to form the node feature matrix X.
[0050] 3) Construct the adjacency matrix: Based on the dependency relationships between nodes, generate the adjacency matrix A, that is, A[i, j]=1 indicates that there is a dependency between nodes v i and v j , and A[i, j]=0 indicates that there is no dependency.
[0051] 4) Data preprocessing: Normalize the node feature matrix X and the adjacency matrix A, including feature standardization and adjacency matrix normalization. Among them, feature standardization is to normalize the node features to ensure that the feature values are in the same range; adjacency matrix normalization is to calculate the symmetric normalized adjacency matrix where D is the degree matrix, representing the degree of each node.
[0052] 5) Data integration: Combine the node feature matrix X and the adjacency matrix A to generate the project network graph data g=(A, X).
[0053] Step 2: Construct the perturbed graph
[0054] Taking the project network graph as the original graph, for each of its nodes, generate a perturbed graph (also known as a perturbed network graph) by removing all the edges related to that node, to simulate the impact on the network structure when a node fails, and retain the node features to ensure that the model can continue to learn the information of that node. The specific processing process is as follows:
[0055] 1) Select the target node: Sequentially select the target node v from the node set V i , for analyzing the importance of this node in the network. Here, each node is perturbed one by one to facilitate quantifying the importance of each node.
[0056] 2) Remove all the edges associated with the target node: Find all the edges in the adjacency matrix A that are connected to v i . A[i, j]=1 or A[j, i]=1 indicates that there is an edge between v i and v j . Set the values of these edges to 0, that is, update the adjacency matrix A to form a perturbed adjacency matrix indicating that all the edges of the target node v i have been removed.
[0057] 3) Retain the node features: After removing all the edges of the target node in the network, the node feature matrix X remains unchanged, and each row in the matrix corresponds to the attribute information of a node, ensuring that the feature X[i] of the target node is retained. This retention method allows the graph neural network to still be able to effectively learn the information of the target node even when the edges are indeed removed, thus ensuring that the key features of the node are still included in the perturbed graph.
[0058] 4) Generate the perturbed graph data: By combining the updated perturbed neighbor matrix with the node feature matrix X, generate the perturbed graph data This data structure reflects the perturbed network topology relationship on the basis of retaining the node features, and serves as the input for subsequent model training, providing data support for node importance evaluation.
[0059] Step 3: Perform negative sampling on the data
[0060] To address the problem of imbalance between positive and negative samples and avoid bias in calculating the loss, negative sampling is performed on the adjacency matrix of the original graph. Edges that do not exist in the network are selected as the negative sample set and participate in model training together with the real edge set, thereby enhancing the model's ability to understand the graph structure. The specific processing process is as follows:
[0061] 1) Extract the node set and relationship set: Extract the node set V and edge set E from the original graph g. Determine the positive sample set M ∈ E, which contains all the actually existing edges e ij whose size is |M| = d i representing the degree of the node v i
[0062] 2) Construct the candidate negative sample set: During the negative sampling process, by traversing the adjacency matrix A, the positions where A ij = 0 in the matrix are used as the basis for the candidate sample set, representing the edges that do not exist in the graph.
[0063] 3) Perform equal sampling: During the negative sampling process, to ensure the balance between positive and negative samples, for each node v i the size of the negative sample set is equal to the size of the positive sample set, that is, the number of negative samples is equal to the number of positive samples. Through random sampling, d i non-existent edges are selected from the candidate negative sample set to form the final negative sample set Thereby ensuring the balance between positive and negative samples during model training and effectively avoiding model bias caused by sample imbalance. The negative sample set includes all e kh satisfying that is, the edges that do not exist in the set E. In this way, a negative sample set can be constructed to ensure that the model can distinguish the non-existent edges in the graph, thereby enhancing the model's ability to understand the graph structure.
[0064] 4) Form the final training samples: Merge the positive sample set M and the negative sample set into the training input and assign a label to each edge. Among them, the label of the positive sample set M is 1, and the label of the negative sample set is 0. This process is used to generate the training data of the model to help the model learn to distinguish positive and negative connections during training, so as to better understand the structural characteristics of the graph.
[0065] Step 4: Reconstruct the perturbed network through a graph neural network
[0066] Construct an encoder f based on a graph neural network θ (GCN, GAT, GraphSAGE, etc.) and decoder p φ , by learning the embedding representation of the perturbed graph , reconstruct the topological structure of the perturbed graph , where the decoder p φ generates the reconstructed adjacency matrix through the embedding vector to obtain the reconstructed graph The specific processing process is as follows:
[0067] 1) Construct the encoder: The role of the encoder f θ is to combine the topological structure of the graph and the node attribute information to generate a low-dimensional, dense embedding representation. To achieve this goal, the present invention can select the following three encoders:
[0068] GCN (Graph Convolutional Network): Use the spectral convolution method to perform embedding learning on the nodes in the graph. The specific form is where is the normalized adjacency matrix, is the degree matrix of the nodes, and local graph structure information can be efficiently captured through this method.
[0069] GAT (Graph Attention Network): Introduce the attention mechanism, which can assign weights to neighbor nodes, enhance the model's attention to important adjacency information, and generate node embeddings through self-attention learning.
[0070] GraphSAGE: Adopt the sampling and aggregation strategy, randomly sample neighbor nodes, and calculate the embedding representation of the target node through the aggregation function, so as to have better computational efficiency when extended to large-scale graph data.
[0071] These encoders extract features from graph data in different ways, providing high-quality embedding representations for the decoder to generate the reconstructed adjacency matrix.
[0072] To meet the requirements of different scenarios in the complex network of engineering projects, the present invention designs a dynamic module selection strategy. Based on network characteristics and task requirements, the encoder is flexibly selected among GCN, GAT, and GraphSAGE to ensure a balance between the performance and efficiency of the graph neural network model.
[0073] ① Input network characteristic analysis: First, pre-analyze the complex network of engineering projects, and extract several key indicators such as the number of nodes, average degree, clustering coefficient, node feature dimension, and global link feature. Among them, the number of nodes measures the network scale, the average degree reflects the network density, the clustering coefficient measures the closeness between nodes, and the node feature dimension characterizes the complexity of node attributes.
[0074] ② Module adaptation rules: Select a suitable encoder according to network characteristics and task requirements. Among them, GCN is suitable for small-scale, dense networks. When the number of network nodes is less than (N < 10 3 ) and the average degree is relatively high, select the GCN module. GAT is suitable for coefficient networks, focusing on key neighbors. When the network is sparse or the task requires highlighting local important neighbors, select GAT; GraphSAGE is suitable for processing large-scale networks. When the network scale is large (N > 10 3 ) or the node feature dimension is high (F > 10 2 ), select GraphSAGE.
[0075] 2) Construct the decoder: The core task of the decoder p φ is to predict the edge connection relationship between nodes in the graph through the node embedding vector, so as to reconstruct the adjacency matrix of the graph. Usually, the decoder consists of two layers of multi-layer perceptrons (MLP). First, perform feature mapping on the input node embeddings, compress the high-dimensional embeddings to an intermediate dimension, and then further map to the final output space to generate a scalar value. The output of the decoder represents the connection probability between node pairs. As a binary classification problem, its output value is between [0, 1], where a value close to 1 indicates that there may be an edge between two nodes, and a value close to 0 indicates that there is no edge. By predicting all possible node pairs through the decoder, a reconstructed adjacency matrix is finally generated. This reconstructed matrix is used to measure the contribution of nodes to the network structure and further guide model optimization. Combining with the loss function, the prediction result of the decoder is compared with the true adjacency matrix A, and the cross-entropy loss is calculated to accurately capture the topological characteristics of the graph structure and the connectivity between nodes.
[0076] Step Five: Node importance calculation
[0077] According to the original graph g=(A, X) and the reconstructed graph Compare the difference between the reconstructed adjacency matrix and the original adjacency matrix, and calculate the cross-entropy loss of node v i to quantify the contribution of the node to the network. This loss can be expressed as:
[0078]
[0079] Since training is based on positive and negative samples, the above loss can be specifically expressed as:
[0080] L = L pos + L eng
[0081] Among them, L pos is the positive sample loss, L negis the negative sample loss. The loss values of positive and negative samples are added together to obtain the total loss value of the current node. The positive sample loss and negative sample loss are expressed as follows:
[0082]
[0083] Among them, M is the set of positive samples (with connected edges); is the set of negative samples (without connected edges); p ij represents the probability that the model predicts the existence of edge e i,j ; w ij represents the weight factor of the edge reconstruction difficulty, which is reflected by the structural entropy H here. The structural entropy H can be used to measure the complexity of nodes and their influence in the network. It considers the local structural characteristics of nodes and can more comprehensively reflect the reconstruction difficulty of nodes. The calculation of w ij is expressed as follows:
[0084]
[0085] Among them, N(i) represents the neighbor set of node v i ; w ik represents the weight of edge e i,k ; if it is an unweighted graph, then w ik = 1; d i is the degree of node v i .
[0086] For each node, repeat the experiment multiple times to obtain a stable loss value, and record its average loss value as the importance index of the node. The specific processing process is as follows:
[0087] 1) Record node loss: Calculate the loss value for each node. The specific operation is as follows: Remove all the edges related to the target node to generate a perturbed graph network. Then, through the graph neural network model, perform embedding learning and network reconstruction on the perturbed graph to obtain the reconstructed adjacency matrix Compare the difference between the reconstructed adjacency matrix and the original matrix A through the cross-entropy loss function. The larger the loss value, the stronger the importance of the node to the network topology and function.
[0088] 2) Obtain stability through multiple experiments: For each node, repeat the above experiment K times (for example, 42 times). Reduce the random noise in a single experiment by randomly initializing the model parameters and randomly negative sampling each time.
[0089] 3) Record the average loss: Take the average of the loss l′ i value of each experiment and record it as the final loss value li i of node v , Specifically expressed as:
[0090]
[0091] Eliminate the randomness influence in the experiment through the average value to ensure that the node importance index has high stability and credibility.
[0092] Step Six: Node Sorting
[0093] Sort the average loss values of all nodes from largest to smallest to obtain the importance sequence L=(l 1 ,l 2 …l N ) of the nodes in the network. The higher the loss value of a node, the greater its impact on the network function and stability, and it needs to be prioritized in resource allocation and risk control. Therefore, these nodes are considered key nodes.
[0094] The processing process of this method (RDGNN) can be represented by the following pseudocode:
[0095]
[0096] Experimental Analysis:
[0097] The present invention verifies the effect on an engineering project network dataset, the structure of which consists of 50 nodes and 222 edges, and the nodes and edges have clear engineering semantics, as Figure 3 shown. Specifically, the nodes represent key entities in the project, such as task execution, resource allocation, and risk management, and their attributes include weight and category. Weight represents the importance or influence degree of the node, such as task priority or resource supply capacity; category classifies the nodes into "task execution" (such as V1, V6), "resource allocation" (such as V3, V9), or "risk management" (such as V2, V8), corresponding to the core components in the project respectively. The edges represent the dependency or interaction relationship between the nodes, such as the collaboration between task nodes and resource nodes, or the sequential dependency between task nodes, reflecting the complexity of multi-dimensional interactions in the engineering project. This dataset truly simulates the relationships among tasks, resources, and risk management in the engineering project, and can be used in research and practical scenarios such as key node identification, resource optimization allocation, and risk assessment, having important engineering application value.
[0098] To verify the effectiveness of the method of the present invention, comparative experiments were carried out with classical key node mining algorithms (such as betweenness centrality, eigenvector centrality, and PageRank). The experimental results are shown in the following table. The top 5 important nodes identified by the PAGNN method are V33, V42, V36, V7, and V41, which are basically consistent with the identification results of other classical methods (such as betweenness centrality, eigenvector centrality, PageRank, etc.), but there are certain differences in the specific importance ranking. Visualize this complex network according to the importance of the nodes. The more important the node is, the closer it is to the center of the network and the denser its connections are. For example, the important nodes V33, V42, and V36 identified by the PAGNN method are located in the core area of the network, have a large number of connections with other nodes, and play a key bridging role. In contrast, less important nodes, such as V17 and V11, are located in the edge area of the network and have fewer connections with other nodes, indicating that they have less impact on the overall structure of the network. In an engineering project network, such highly important nodes located in the center (such as V33 and V42) often represent key roles in resource scheduling or task allocation, and their failure may have a significant impact on the stability of the entire project, so they need to be given priority attention.
[0099] Table 1 Top 15 Nodes of the Complex Network of Engineering Projects
[0100] Rank PAGNN Betweenness centrality Eigenvector centrality PageRank 1 V33 V33 V33 V33 2 V42 V12 V7 V42 3 V12 V36 V42 V36 4 V36 V42 V36 V12 5 V7 V41 V12 V48 6 V48 V48 V48 V7 7 V41 V7 V25 V41 8 V25 V25 V41 V20 9 V29 V20 V29 V25 10 V20 V6 V20 V29 11 V11 V30 V30 V30 12 V44 V44 V11 V44 13 V6 V29 V44 V6 14 V30 V17 V6 V17 15 V17 V11 V17 V11
[0101] Although the present invention has been disclosed above in embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall be covered within the protection scope of the present invention. The protection scope of the present invention shall be defined by the claims.
Claims
1. A key node mining method for complex networks of engineering projects, characterized in that: The following steps are involved: 1) Construct a project network graph based on the engineering project network and use it as the original graph. The nodes in the graph represent project entities, the edges represent the dependencies between projects, and define the node feature matrix; 2) Generate a perturbation graph by removing all edges of the target node in the original graph and retaining the node feature matrix; 3) Perform data sampling on the original image to obtain a set of positive and negative samples and construct a training data set; 4) The perturbation graph is reconstructed using the encoder and decoder of the graph neural network, where the encoder generates a node embedding vector, the decoder reconstructs the adjacency matrix, and the reconstruction capability of the decoder is trained using the above training data set; 5) Compare the differences between the original graph and the reconstructed graph and calculate the loss value of the node; 6) Sort the nodes by importance according to their loss values and determine the key nodes.
2. The method according to claim 1, characterized in that The steps of constructing a project network diagram according to the engineering project network in step 1) include: Represent the project network as an undirected and unweighted graph, and define the node set, edge set and node feature matrix; According to the attribute information of the node, the node features are extracted and the node feature matrix is constructed; Generate an adjacency matrix based on the dependencies between nodes; Normalize the node feature matrix and adjacency matrix; Combine the node feature matrix and the adjacency matrix to generate project network graph data.
3. The method according to claim 2, characterized in that Node characteristics include resource consumption, project contribution and project cost-effectiveness.
4. The method according to claim 2, characterized in that Normalizing the node feature matrix includes normalizing the node features, and normalizing the adjacency matrix includes calculating a symmetric normalized adjacency matrix.
5. The method according to claim 1, characterized in that The step of generating a perturbation graph in step 2) further includes: After removing all edges associated with the target node, update the adjacency matrix to generate a perturbed adjacency matrix; The perturbed adjacency matrix is combined with the node feature matrix to generate perturbed graph data.
6. The method according to claim 1, characterized in that The step of sampling data on the original graph in step 3) includes: extracting a node set and an edge set from the original graph, and constructing a positive sample set according to the actually existing edges; Select the edges that do not exist in the adjacency matrix as candidate negative samples and construct a set of candidate negative samples; From the candidate negative sample set, non-existent edges are selected by random sampling to form a negative sample set, and the size of the negative sample set is equal to that of the positive sample set; Merge the positive sample set and the negative sample set, and assign a label to each edge to generate training data.
7. The method according to claim 1, characterized in that In step 4), the encoder dynamically selects GCN, GAT or GraphSAGE according to the network characteristics and task requirements of the engineering project.
8. The method according to claim 1, characterized in that The step of reconstructing the adjacency matrix by the decoder in step 4) includes: Through two multi-layer perceptrons, the input node embedding is first feature mapped to compress the high-dimensional embedding to the intermediate dimension; then the intermediate dimension is mapped to the final output space to generate a scalar value, which represents the connection probability between nodes, predicts the edge relationship between nodes, and generates a reconstructed adjacency matrix.
9. The method according to claim 1, characterized in that The node loss value calculated in step 5) is specifically the mean cross entropy loss of each node in multiple tests, which is composed of the sum of the mean positive sample loss and the mean negative sample loss.
10. A key node mining system for complex networks of engineering projects, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 9 when executing the computer program.