Complex network key node identification method based on graph convolutional neural network

By using graph convolutional neural networks in complex networks, combining local, community and global information, and adopting multi-scale feature fusion and channel attention mechanisms, the problems of low recognition accuracy and poor adaptability in the existing technology are solved, and efficient identification of key nodes of complex networks are achieved.

CN120046016APending Publication Date: 2025-05-27CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118762.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing methods for identifying key nodes in complex networks usually only consider single aspects of information, resulting in low recognition accuracy and difficulty in adapting to different network structures, especially in sparse or disconnected networks.

Method used

A method for identifying key nodes in complex networks based on graph convolutional neural networks is proposed. By obtaining local, community and global information of each node and inputting it into a multi-scale convolutional neural network, the node importance score is automatically learned by utilizing multi-scale feature fusion and channel attention mechanism.

Benefits of technology

It realizes efficient identification of key nodes in complex networks and improves recognition accuracy, especially in sparse or disconnected networks, which outperforms traditional methods and other deep learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046016A_ABST
    Figure CN120046016A_ABST
Patent Text Reader

Abstract

The invention provides a complex network key node identification method based on a graph convolutional neural network. The method comprises the following steps: S1, obtaining a neighborhood network of each node; s2, obtaining local, community and global information according to the network; s3, inputting the data information into a multi-scale network; and S4, outputting the identified key nodes. The key nodes are identified according to local, community and global information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of complex networks, and in particular to a method for identifying key nodes of complex networks based on graph convolutional neural networks. Background Art

[0002] Nowadays, many real-world systems can be modeled by complex networks. Among numerous complex networks, a few nodes often have a profound impact on the structure and dynamic processes of the entire network. Therefore, identifying these key nodes has become an important task in complex network analysis and plays an important role in multiple practical application scenarios, such as epidemic spread, social network analysis, power network protection, and new drug research and development. Therefore, the identification of key nodes has become a research hotspot in the field of complex networks. For example, in gene networks, identifying key genes related to the occurrence of diseases can accelerate the treatment process; in traffic networks, regularly maintaining key hub roads can effectively prevent the collapse of the traffic system; in social networks, it helps to control the spread of information and rumors and identify leaders in the network. Finding bloggers who have a significant influence on the target customers of a specific company helps to promote new products more efficiently. Summary of the Invention

[0003] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a method for identifying key nodes of complex networks based on graph convolutional neural networks.

[0004] To achieve the above object of the present invention, the present invention provides a method for identifying key nodes of complex networks based on graph convolutional neural networks, including the following steps:

[0005] S1, obtaining the neighborhood network of each node; the node can include one of a shopping node, a product node, etc.;

[0006] S2, obtaining local, community, and global information according to the network;

[0007] S3, inputting the data information into a multi-scale network;

[0008] S4, outputting the identified key nodes.

[0009] In a preferred embodiment of the present invention, in step S2, the generation method of local information is:

[0010]

[0011] where SLC(i) represents the value of the semi-local centrality of node i;

[0012] Γ(i) is the set of direct neighbors of node i;

[0013] Γ(j) is the set of direct neighbors of node j;

[0014] Q(j) represents the sum of the degrees of the second - order neighbors of node j;

[0015] N(w) is the number of second - order neighbors of node w;

[0016]

[0017] Among them, X Loc represents the local hierarchical information matrix of the node;

[0018] A ij represents the adjacency matrix of the network;

[0019] SLC(j) represents the value of the semi - local centrality of node j;

[0020] r represents the number of neighbors of a node (when the total number of neighbors of the node < neighborhood size S);

[0021] md is the neighbor aggregation degree of the corresponding node;

[0022] S represents the neighborhood size;

[0023] SLC(i) represents the value of the semi - local centrality of node i;

[0024] In a preferred embodiment of the present invention, in step S2, the method for generating community information is: VC(i) = |V(Γ(i))|,

[0025] where VC(i) represents the importance value of node i in the community;

[0026] V(Γ(i)) represents the set of communities to which node i and its neighbors belong;

[0027] |V(Γ(i)| represents the number of the set of communities to which node i and its neighbors belong;

[0028]

[0029] Among them, X Com represents the community hierarchical information matrix of the node;

[0030] A ij represents the adjacency matrix of the network;

[0031] VC(j) represents the importance value of node j in the community;

[0032] r represents the number of neighbors of a node (when the total number of neighbors of the node < neighborhood size S);

[0033] md is the neighbor aggregation degree of the corresponding node;

[0034] S represents the neighborhood size;

[0035] VC(i) represents the value of the importance of node i in the community.

[0036] In a preferred embodiment of the present invention, in step S2, the method for generating global information is as follows:

[0037]

[0038] Among them, EC(i) represents the value of the eigenvector centrality of node i;

[0039] x i is the eigenvector centrality value of node i;

[0040] c is a proportionality constant;

[0041] n represents the number of nodes;

[0042] a ij represents the element in the i-th row and j-th column of the adjacency matrix;

[0043] x j represents the eigenvector centrality value of node j.

[0044]

[0045] Among them, X Glo represents the entire hierarchical information matrix of the nodes;

[0046] A ij represents the adjacency matrix of the network;

[0047] EC(j) represents the value of the eigenvector centrality of node j;

[0048] r represents the number of neighbors of a node (when the total number of neighbors of the node < S);

[0049] md is the neighbor aggregation degree of the corresponding node;

[0050] S represents the neighborhood size;

[0051] EC(i) represents the value of the eigenvector centrality of node i.

[0052] In a preferred embodiment of the present invention, it further includes:

[0053]

[0054] Among them, DC(i) represents the value of the degree centrality of node i;

[0055] k i is the degree of node i;

[0056] n is the number of nodes.

[0057] In a preferred embodiment of the present invention, the multi-scale network consists of three convolutional layers, three pooling layers and one fully-connected layer.

[0058] In a preferred embodiment of the present invention, after step S4, step S5 is further included to verify the indicators of the multi-scale network.

[0059] In a preferred embodiment of the present invention, the indicators include one or any combination of a ranking measurement indicator, a monotonicity index of the ranking result, and an overlap degree indicator.

[0060] In a preferred embodiment of the present invention, the calculation method of the ranking measurement indicator is:

[0061]

[0062] where τ represents the ranking measurement indicator;

[0063] N c represents the number of node consistent pairs;

[0064] N d represents the number of node inconsistent pairs;

[0065] N L represents the number of nodes in the sorted list.

[0066] In a preferred embodiment of the present invention, the calculation method of the monotonicity index of the ranking result is:

[0067]

[0068] where MI represents the monotonicity index of the ranking result;

[0069] L 0 is the ranking value list;

[0070] N α is the number of nodes with the same ranking value α;

[0071] N L is the number of nodes in the ranking list;

[0072] In a preferred embodiment of the present invention, the calculation method of the overlap degree indicator is:

[0073]

[0074] where OVER-K represents an indicator measuring the overlap degree between the key nodes selected by the algorithm and the actually important nodes;

[0075] is the set of the top k importance prediction nodes;

[0076] ∩ represents the intersection;

[0077] S top-k is the set of the top k true importance nodes;

[0078] represents the intersection of the set of the top k importance prediction nodes and the set of the top k true importance nodes.

[0079] The present invention also discloses a computer system, including:

[0080] a processor;

[0081] a memory for storing instructions executable by the processor;

[0082] wherein, when the processor is configured to execute the executable instructions, the method for identifying key nodes of a complex network based on a graph convolutional neural network is implemented.

[0083] The present invention also discloses a computer-readable storage medium, including:

[0084] a memory having a computer program stored thereon;

[0085] a processor for executing the program in the memory to implement the method for identifying key nodes of a complex network based on a graph convolutional neural network.

[0086] In summary, due to the adoption of the above technical solutions, the present invention realizes the identification of key nodes according to local, community, and global information, and realizes product recommendation.

[0087] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, wherein:

[0089] Figure 1 is a schematic diagram of the MSACNN algorithm framework of the present invention.

[0090] Figure 2 is a schematic diagram of the sorting ability of the nodes of the present invention.

[0091] Figure 3 is a schematic diagram of the influence of the size of the neighborhood network on the algorithm of the present invention.

[0092] Figure 4It is a schematic diagram of the experimental results of different component combinations in the algorithm adopted by the present invention.

[0093] Figure 5 It is a schematic diagram of the correlation of the propagation influence score of the present invention. Detailed implementation manners

[0094] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0095] Regarding the problem that the traditional key node recognition method for complex networks generally starts from the network topology properties and either considers a single index or manually assigns weights to multiple indexes, resulting in low recognition accuracy, etc.; the present invention's patent proposes an improved graph convolutional neural network key node recognition method (Multi-Scale-Attention-CNN, MSACNN) that comprehensively considers local, community, and global three-dimensional structure information; first, introduce the breadth-first algorithm to extract a neighborhood network of size S for each node, and extract the information of the low-complexity semi-local centrality algorithm (SLC), the number of communities to which the node is connected, and the eigenvector centrality algorithm (EC) and embed it into the adjacency matrix of the node; secondly, propose a way to aggregate the neighbor degrees to solve the problem of sparse feature matrices on disconnected networks; input the feature matrix into a multi-scale convolutional neural network (multi-scale network) that fuses the channel attention mechanism to extract node information, and the features of the three structural attributes can be automatically learned during the training process to obtain the node importance score; finally, conduct a comparative analysis experiment on 8 datasets such as Arenas and 7 baseline methods such as GAT. The experimental results show that: the proposed MSACNN algorithm improves by 0.2% to 10.07% compared with the baseline methods in three indexes such as Kendall’s τ, indicating the rationality and effectiveness of the MSACNN method.

[0096] Keywords: complex network; key node recognition; graph convolutional network; neighborhood network

[0097] 1 Introduction

[0098] 1.1 Background

[0099] Node ranking is an important topic in the study of complex networks, and many scholars have proposed different node ranking methods. The main methods for evaluating the importance of nodes in complex networks can be divided into three categories: traditional methods, machine learning-based methods, and deep learning-based methods. Traditional methods can be further subdivided into two categories. One category is the measurement methods based on network topology. These methods usually evaluate the importance of nodes by analyzing the local information of nodes, such as degree centrality, betweenness centrality, K-shell algorithm, etc. These methods judge the relative importance of nodes by considering their positions in the network. Specifically, DC defines the number of first-order neighbors of a node as the importance degree of that node. BC defines the nodes that act as bridges between two disconnected groups as nodes that spread influence. The focus of K-shell decomposition is the position of nodes in the network. The above sorting algorithms based on deterministic metrics and rules perform well in specific networks. However, their relatively limited expressive power makes them unable to be widely applied to multiple networks. In addition, these methods ignore the characteristic information of the nodes themselves. Another category of traditional methods is based on the destructiveness of the network. This idea stems from the fact that node deletion may weaken the robustness of the network, such as cascading failures. In this case, the greater the influence of a node, the worse the robustness of the network. Different from traditional physical methods, machine learning models can automatically learn and optimize the weights of different attributes during the training process. Based on traditional physical methods that only rely on microstructure information, machine learning-based methods show stronger generalization ability by integrating multiple structural attributes. Specifically, machine learning-based methods can be further divided into algorithms based on statistical machine learning and algorithms based on deep learning. Machine learning-based methods for evaluating node importance include gradient boosting decision tree (GBDT), least squares support vector machine (LS-SVM), and ensemble learning (EL), etc. These methods can more accurately evaluate the importance of nodes by learning the patterns in the network.

[0100] Deep learning-based node importance evaluation methods include Convolutional Neural Network (CNN), Graph Neural Network (GNN), Reinforcement Learning (RL), etc. These methods lay a solid foundation for further research on node importance ranking. Compared with traditional machine learning methods, deep learning methods reduce the dependence on community detection algorithms by introducing richer node structure information, thus being able to capture key features in complex networks more precisely. Deep learning-based algorithms are mainly based on Graph Convolutional Network (GCN) because they can automatically filter important features. However, with more and more structural attributes being considered, how to balance efficiency and accuracy in designing machine learning-based algorithms has become a new challenge. Since it can automatically screen out important features, they show strong advantages in dealing with complex networks. However, with the introduction of more structural attributes, how to balance efficiency and accuracy in designing machine learning-based algorithms has become a new challenge.

[0101] Research shows that the propagation influence of a node is closely related to the structural characteristics of its affiliated community. The characteristics of the community structure enable the nodes in the network to be divided into multiple tightly connected node sets, and the nodes within each community are relatively closely connected. The number of communities a node connects to can be used as an indicator to identify propagation influence nodes that may be overlooked by centrality metrics. However, when using different community detection methods, the number of communities in the network may change, which challenges the stability of this strategy. To overcome this problem, Zhang et al. proposed a meta-heuristic-based community detection search method. In the agglomerative clustering process of this method, each unit is regarded as an independent community in the initial stage; subsequently, the communities will be continuously merged until the specified performance metric cannot be further improved.

[0102] However, there are still some deficiencies in the existing research. First, algorithms based on heuristic rules are usually limited by their expressiveness, so their performance may vary greatly in different application scenarios. Recently, Yu et al. proposed an efficient algorithm based on graph convolutional network, called the RCNN algorithm. This algorithm extracts a fixed-size neighborhood network for each node through breadth-first search (BFS), and uses the adjacency matrix and degree value of each node's neighborhood network as the input of the convolutional neural network (CNN). Although the RCNN algorithm can adapt to large-scale networks, it only relies on the degree value of nodes to construct the input, which may lead to nodes located at the edge of the network being misjudged as important propagation nodes. To solve this problem, Ouyang proposed an improved multi-channel M-RCNN algorithm. This algorithm takes into account the structural information at the micro level, community level, and macro level. However, when considering the macro-level information, this algorithm uses the k-shell algorithm, and the BA network used in training the model will make the k-shell algorithm ineffective. As a result, the training of the model is not complete enough, and there are obvious disadvantages in the application of disconnected networks. This invention patent proposes a multi-scale convolutional neural network (MSACNN) based on GCN, uses a more efficient algorithm to obtain the input features of the convolutional network, and adds a multi-scale convolutional feature fusion attention mechanism to improve the performance of the model. The proposed method has excellent performance on nine real network datasets.

[0103] 1.2 Research Motivation

[0104] In the research based on network science in the past few decades, finding key nodes in complex networks has always been an enduring research hotspot. Currently, there are many methods to identify key nodes. However, the centrality method based on network structure is limited by the limitations of manually selected features and often only considers information from a single aspect. In the recently emerging methods based on deep learning, how to obtain efficient node representations and design neural networks has also become a difficult point. It is very important to find an efficient method for node representation in dealing with special networks such as sparse or disconnected networks.

[0105] 1.3 Main Contributions

[0106] (1) To effectively extract the local, community, and global information of nodes, a method for extracting three-dimensional complex network node features is designed; a method for aggregating the degrees of node neighbors is proposed to make up for the problem of sparse feature matrices caused by insufficient subgraph sizes in disconnected complex networks.

[0107] (2) To process the multi-dimensional information of nodes, a weighted fusion multi-scale convolutional neural network is designed to perform weighted fusion on different dimensions of information on the feature matrices extracted by convolutional kernels of different sizes and automatically train the weights.

[0108] (3) A method for capturing features with a channel attention mechanism on a multi-scale convolutional neural network is proposed, which dynamically adjusts the contribution of each channel to the final output, thus more effectively capturing the features most crucial for the current task and suppressing irrelevant or redundant information.

[0109] (4) Considering the local, community, and global three-dimensional structure information comprehensively, a method for identifying key nodes in a multi-scale graph convolutional neural network with an attention mechanism (MSACNN) is proposed.

[0110] 1.4 Related Work

[0111] 1.4.1 Traditional Methods

[0112] Traditional methods have been widely studied in recent years and can be mainly divided into four categories of network metric methods. (1) Global structure metric methods: These methods include betweenness centrality, eigenvector centrality, etc. Fei proposed a ranking method based on node strength, which depends on the distance between nodes; Lv proposed a method considering the change of the average shortest path in the network. However, the computational complexity of these methods is relatively high, which limits their application in large-scale networks. (2) Local structure metric methods: These methods mainly include degree centrality, closeness centrality, etc. For example, Xu proposed an adjacent information entropy method based on the calculation of node adjacency degree; Sheng proposed an algorithm combining closeness centrality and nearest neighbor nodes. The main disadvantage of such methods is that they only focus on the local information of nodes and ignore the environment and position of nodes in the whole network. (3) Network dynamic characteristic analysis methods: For example, Google's PageRank algorithm and VoteRank algorithm, etc. Liu et al. combined eigenvector centrality (EC) with the gravity model and proposed the concept of weighted gravity centrality. However, there are inconsistencies between the path information obtained by eigenvector centrality and the local information provided by the gravity model. (4) Node localization methods: These methods mainly include the K-shell algorithm, etc. The local degree dimension (LDD) takes the number of neighbors of the central node in each layer as an important indicator to measure the importance of nodes. The limitation of these methods is that the results are relatively rough, and some nodes may have the same K-shell value, resulting in insufficient discrimination between nodes and making it difficult to accurately evaluate the importance of nodes.

[0113] 1.4.2 Machine Learning and Deep Learning-Based Methods

[0114] In recent years, with the wide application of machine learning methods in multiple fields, researchers have gradually introduced machine learning into the node ranking problem in complex networks, mainly by feature selection or feature engineering to improve performance. A variety of machine learning techniques, such as logistic regression, support vector machine (SVM), and K-nearest neighbor (KNN), have been widely used to identify key users. Wen used LS-SVM to train an evaluation model and obtained evaluation results. Fan found an optimal node set through the RL algorithm, where the states and actions of the graph were represented by inductive graph representation learning. Li used the EL algorithm to identify necessary nodes by estimating network robustness. These techniques have been applied in various industries and different scenarios, laying the foundation for many advances in this field. However, a major drawback of using these algorithms is that the feature selection process is very time-consuming, and at the same time, they often ignore the connection relationships between target nodes.

[0115] As a research field that has received much attention, a variety of methods based on graph convolutional networks (GCNs) have been developed to identify key nodes in complex networks. These methods not only solve the problem of node relative importance but also address a variety of other challenges. GCN (graph convolutional network) is a semi-supervised learning method for graph data, which can usually be trained by backpropagation and task-specific loss functions like a standard convolutional neural network (CNN). However, due to the lack of a grid structure in graph data, standard convolutional operations cannot be directly applied to image or text data, so the core difficulty of GCN lies in how to effectively extract structural information from the graph. Zhao et al. first integrated the local structural information of nodes through the superposition of neighbor graphs and used four classic structural features as the node representation for the input of the graph convolutional network. Qu et al. proposed a time information gathering (TIG) process to evaluate the importance of nodes in a temporal network. In the TIG process, the importance of a node depends on the importance of its neighborhood. Cai et al. proposed an efficient and direct method to improve the graph contrast learning framework in recommendation systems. Specifically, they improved the structure of the user-item interaction graph by enhancing the ability of singular value decomposition. The research results show that the proposed graph enhancement method can effectively alleviate the problems of data sparsity and popularity bias. Wei et al. proposed a novel contrastive graph structure learning method based on information bottleneck (CGI), aiming to learn improved multi-view representations from different perspectives. They proposed a fully differentiable learner to remove nodes and edges and created various forms of enhanced views by combining with recommendation systems. By innovatively introducing the information bottleneck into multi-view contrastive learning, the authors verified the effectiveness of this method. Recently, Yu transformed the significant node identification problem into a regression problem through CNN. The labels were determined by the infection scale, and the feature matrix was obtained through the adjacency matrix and node degree. OU et al. proposed a new application based on GCN to identify influential nodes through multi-level structural attributes.

[0116] 2Preliminary

[0117] 2.1 Definitions

[0118] This invention patent represents a complex network as a graph G = (V, E), which consists of |V| = n nodes and |E| = m edges. Among them, V = {v 1 , v 2 , …, v n} is the node set, and E = {e 1 , e 2 , …, e m} is the edge set. Among them, G represents an undirected network (graph), V represents the node set, E represents the edge set, |V| represents the number of nodes in the node set V, and its value is n; |E| represents the number of edges in the edge set E, and its value is m; v 1 represents the first node, v 2 represents the second node, v n represents the nth node, e 1 represents the first edge, e 2 represents the second edge, e m represents the mth edge; let n be the number of nodes and m be the number of edges. Each network can be represented by an adjacency matrix A ij of size n×n. Table 1 introduces the mathematical symbols used in this invention patent.

[0119] Table 1 Introduction to Mathematical Symbols

[0120]

[0121] 2.2 Assumptions

[0122] The network is a static undirected network. To simplify the discussion and focus on the core idea of the proposed algorithm, this study assumes that the network is a static undirected graph. Although the algorithm can be extended to support various types of networks (including directed networks and time-varying networks), covering all these cases in a single study would add unnecessary complexity. By choosing a static undirected network as the basic model, we can provide a clear and easy-to-understand framework, which not only helps readers better master the core concept of the algorithm but also lays a solid foundation for future research, facilitating subsequent extension to more complex network types.

[0123] Assume that the network has a clear community structure. Without considering overlapping communities, the Louvain algorithm is adopted for community detection in this study. The Louvain algorithm identifies the optimal community partition by optimizing modularity, but it does not consider the situation where a node may belong to multiple communities simultaneously. Therefore, when training and testing the network, we assume that each node belongs to only one specific community to ensure the clarity of community boundaries. This assumption helps to analyze the dynamics within and between communities more deeply and provides support for the effective implementation of the algorithm.

[0124] 2.3 The proposed method

[0125] The research of this invention patent aims to solve the problem of identifying key nodes in complex networks through a multi-scale convolutional neural network (MSACNN), and proposes a representation learning framework that can handle arbitrary network structures. First, the BFS algorithm is used to obtain network subgraphs, and the most relevant nodes are preferentially selected according to the hop count and degree of neighbors. Through these neighborhoods, we construct a local subnetwork that contains the structural information of nodes and their neighbors. Further, the local, community, and global information of the node is embedded into the adjacency matrix of the subnetwork as the input of the multi-scale attention convolutional neural network, and the infection scale on the SIR propagation model is used as the label to predict the node infection scale. Finally, the effective representation of the network is learned through this framework to identify the key nodes in the network. The schematic diagram of the proposed algorithm framework is as Figure 1 shown.

[0126] Figure 1 is the schematic diagram of the MSACNN algorithm framework; among them, (a) is to extract nodes and edges in a real network; (b) is to generate the adjacency matrix of the neighborhood network; (c) is to convert the adjacency matrix into a feature matrix according to the rules; (d) is the multi-scale convolutional layer and pooling layer; (e) is the channel attention layer, which learns channel weights to capture the most effective features; (f) is the fully connected layer that outputs importance scores and trains according to the label; (g) is to obtain the sorting result according to the node importance scores.

[0127] 1) Feature extraction

[0128] First, the BFS algorithm is used to extract node features. Specify S as the size of the neighborhood network for each node in the network, and extract the neighborhood network for each node. The method algorithm is as follows:

[0129]

[0130]

[0131] Stop when the number of nodes added to the neighborhood network accumulates to S - 1. All these nodes will be encoded according to the order in which they are incorporated. During this process, neighbor nodes determine the order of joining the neighborhood network based on their distance from the target node and their own degree; that is, nodes closer to the target node and with higher degrees are preferentially selected. If there is a situation where two neighbor nodes are equidistant from the target node and have the same degree, their joining order will be randomly determined. Finally, an adjacency matrix A representing this neighborhood network is obtained. ij , which forms the basis of node representation. It should be noted that if the total number of direct or indirect neighbors of the target node is less than S - 1, we first fill the feature matrix with zeros, and then use the average neighbor degree to fill this part of the information, ensuring that the obtained adjacency matrix reaches the expected scale and effectively improving the performance of the algorithm on sparse graphs.

[0132] 2) Feature transformation

[0133] After obtaining the adjacency matrix and considering the problem of feature engineering, we construct a three-channel input for the multi-scale convolutional neural network (multi-scale network). Here, three algorithms with lower complexity and wide coverage of information are used to obtain the information of nodes at the local, community, and global levels respectively. At the local level, the SLC algorithm is used to obtain information because it can obtain information of higher-order neighbors, and at the same time, the first-order and second-order neighbor numbers of the direct neighbors of the node are considered; at most, the information of the fourth-order neighbors of the node is involved. The calculation formula of SLC is:

[0134]

[0135] Among them, SLC(i) represents the value of the semi-local centrality of node i;

[0136] Γ(i) is the set of direct neighbors of node i;

[0137] Γ(j) is the set of direct neighbors of node j;

[0138] Q(j) represents the sum of the second-order neighbor degrees of node j;

[0139] N(w) is the number of second-order neighbors of node w;

[0140] Generate the local information representation matrix of the node according to the following rules:

[0141]

[0142] Among them, X Loc represents the local-level information matrix of the node;

[0143] A ij represents the adjacency matrix of the network;

[0144] SLC(j) represents the value of the semi-local centrality of node j;

[0145] r represents the number of neighbors of a node (when the total number of neighbors of the node < neighborhood size S);

[0146] md is the neighbor aggregation degree of the corresponding node;

[0147] S represents the neighborhood size;

[0148] SLC(i) represents the value of the semi-local centrality of node i;

[0149] At the community level, the louvain algorithm is used to give information about the node community. The more communities a node is adjacent to, the more important the node is. VC represents the importance of the node, and the calculation formula of VC is:

[0150] VC(i) = |V(Γ(i))| (3)

[0151] Among them, VC(i) represents the importance value of node i in the community;

[0152] V(Γ(i)) represents the set of communities to which node i and its neighbors belong;

[0153] V(Γ(i)| represents the number of the set of communities to which node i and its neighbors belong;

[0154] Generate the community information representation matrix of the node through the following rules:

[0155]

[0156] Among them, X Com represents the community level information matrix of the node;

[0157] A ij represents the adjacency matrix of the network;

[0158] VC(j) represents the importance value of node j in the community;

[0159] r represents the number of neighbors of a node (when the total number of neighbors of the node < neighborhood size S);

[0160] md is the neighbor aggregation degree of the corresponding node;

[0161] S represents the neighborhood size;

[0162] VC(i) represents the value of the importance of node i in the community;

[0163] At the global level, the eigenvector centrality algorithm is used to obtain information. Since the commonly used K-shell algorithm fails in training the BA network, and the eigenvector centrality (EC) can well reflect the global importance information of nodes, the calculation formula of EC is as follows:

[0164]

[0165] Among them, EC(i) represents the value of the eigenvector centrality of node i;

[0166] x i is the eigenvector centrality value of node i;

[0167] c is a proportionality constant;

[0168] n represents the number of nodes;

[0169] a ij represents the element in the i-th row and j-th column of the adjacency matrix;

[0170] x j represents the eigenvector centrality value of node j.

[0171] The global information representation matrix of nodes is generated through the following rules:

[0172]

[0173] Among them, X Glo represents the entire hierarchical information matrix of nodes;

[0174] A ij represents the adjacency matrix of the network;

[0175] EC(j) represents the value of the eigenvector centrality of node j;

[0176] r represents the number of neighbors of a node (when the total number of neighbors of the node < S);

[0177] md is the neighbor aggregation degree of the corresponding node;

[0178] S represents the neighborhood size;

[0179] EC(i) represents the value of the eigenvector centrality of node i;

[0180] By the above method, the information of the three levels is fused into the adjacency matrix A ij to construct three new feature matrices X Loc , X Com , X Glo to be used as the input of the multi-scale convolutional neural network.

[0181] 3) Obtain labels

[0182] The SIR model is a classic infectious disease transmission model and is also widely used in the study of information diffusion. This model divides the individuals in the network into three states: Susceptible, Infected, and Recovered. In the initial state, except for the specified source node being in the infected state, all other nodes are in the susceptible state. As time progresses, at each time step, an infected node will attempt to spread its state to adjacent susceptible nodes with a certain probability. Once successful, these susceptible nodes will change to the infected state. In addition, an infected node will also convert to the recovered state with a fixed recovery probability, and nodes entering this state will acquire immunity and no longer be infected. This process continues until no new infections occur, that is, all nodes that could potentially be infected have either experienced infection or remained uninfected. Finally, the total number of nodes in the recovered state can be used as an indicator to measure the influence of the source node.

[0183] In the SIR model, choosing an appropriate transmission probability is crucial for accurately evaluating the influence of nodes. Our research focuses on the transmission probability near the epidemic threshold. Specifically, the formula for calculating the critical infection probability of a real network is u c = <k> / (<k 2 >- <k>), where <k>Denotes the average degree of nodes, <k2>is the second moment of the node degrees; and for a random network, the critical propagation probability simplifies to 1 / <k>。The recovery rate set here is 1. To ensure the reliability of the results, we performed 1000 independent runs for each configuration and retained the average of these run results as the label.

[0184] 4) Model training

[0185] After preparing the input and labels of the nodes, the network is trained. This invention patent constructs a multi-scale convolutional feature fusion attention neural network (multi-scale network), which consists of three convolutional layers, three pooling layers, and one fully connected layer. Specifically, three different-sized convolutional kernels of 3*3, 5*5, and 7*7 are used in the convolutional layers to obtain information of different dimensions and perform weighted fusion. The input channels of the first convolutional layer are 3, and the output channels are 10. Before inputting the first channel, we also assign learnable weights to each channel to adaptively adjust the importance of three-dimensional information. The input channels of the second convolutional layer are 10, and the output channels are 20. The input channels of the third convolutional layer are 20, and the output channels are 30. The stride is 1, and the paddings are 1, 2, and 3 respectively. The activation function is leakyrelu. The pooling layer uses 2*2 max pooling. After three times of convolution and pooling, the channel attention mechanism is applied. First, by learning the importance weights of each channel, the contribution of different channels to the final output is dynamically adjusted, so that the model can more effectively capture the features most useful for the current task. Finally, the captured features are integrated through a fully connected layer to output the importance scores of the nodes. The loss function uses MSE (mean square loss function). For the training of the model, BA scale-free networks of different sizes are used as the training set, and the corresponding SIR infection scale ranking results are used as the labels. The trained model is applied to the feature matrices generated by different unlabeled real networks to obtain the importance scores of all nodes and sort them. Comparing the obtained results with the real SIR infection ranking results, it is found that the proposed MSACNN method has excellent performance on different real networks.

[0186] 5) Complexity analysis

[0187] Analyzing the complexity of the MSACNN algorithm proves that the selected feature extraction algorithm is simple and effective. The computational complexity of the SLC algorithm is O(n <k> 2 ), the computational complexity of VC is O(nlogn), and the computational complexity of eigenvector centrality is O(m + n). In addition, the computational complexity of the BFS algorithm is O(n + m). Therefore, the computational complexity of the input construction step of the proposed algorithm is O(n <k> 2 +n log n + n + m + n + m) ≈ O(m + n log n). The computational complexity of the SIR model in the worst case is O(Tn), where T represents the number of extended time steps. The computational complexity of the three multi-scale convolutional layers is where C in and C out represent the number of input channels and output channels respectively, K represents the convolutional kernel size, and the complexity of weighted fusion is O(C out × S 2 ). S represents the size of the input feature map, the complexity of the pooling layer is O(S 2 ), and the computational complexity of the fully connected network is O(C out × (S / 8) 2 ) = O(C out × S 2 / 64), and n represents the number of nodes; <k> 2 represents the square of the average degree; m represents the number of edges; K 1 represents the size of the first convolutional kernel; K 2 represents the size of the second convolutional kernel; K 3 represents the size of the third convolutional kernel. It is proved that the proposed MSACNN has a lower complexity. It is worth noting that the experimental results show that the proposed algorithm trained on a small-scale network such as a BA network with 1000 nodes and an average degree of 4 can be used to predict the diffusion influence of nodes in a large-scale network.

[0188] 3 Experimental Results and Analysis

[0189] Kendall's τ, MI, and OVER-K were used as three evaluation metrics to conduct comparative analysis experiments with 7 baseline methods on 9 different real network datasets.

[0190] 3.1 Datasets

[0191] To verify the effectiveness and discriminability of the MSACNN model, 9 different real networks were selected from different fields such as communication, power grid, social interaction, biology, and music for experiments in this section. They have different topological characteristics. The training network is a BA network with 1000 nodes and an average degree of 4.

[0192] The BA network is a scale-free network with n nodes and an average degree of k generated by the BA model.

[0193] Arenas: The metabolic network of Caenorhabditis elegans.

[0194] GrQC: The author research collaboration network in the field of general relativity and quantum cosmology, where nodes are authors and an edge is established when two authors co-author a paper.

[0195] Pgp: The dataset shows a user network called Pretty-Good-Privacy for secure data exchange algorithms

[0196] Hep: The author research collaboration network in the field of high-energy physics, where nodes are authors and an edge is established when two authors co-author a paper.

[0197] Stelzl: The network represents pairs of interacting proteins in humans (Homo sapiens).

[0198] Figeys: The network of interactions between human proteins, from the first large-scale study of protein-protein interactions in human cells using a mass spectrometry-based method.

[0199] jazz: The collaboration network between jazz musicians.

[0200] Powergrid: A network containing power grid information of the western states of the United States.

[0201] NS: A collaborative network of scientists dedicated to network theory and experiments.

[0202] Table 2 lists some statistical properties of the real networks. The networks used in this invention patent are undirected and unweighted networks.

[0203] Table 2 Structural information and related statistical data of nine datasets

[0204]

[0205]

[0206] 3.2 Baseline methods

[0207] This invention patent uses the following methods for comparison, and these methods can be roughly divided into two categories of algorithms. Four centrality algorithms based on network structure and three deep learning-based algorithms.

[0208] DC: Degree centrality defines the number of first-order neighbors of a node as the importance degree of that node, and its definition is:

[0209]

[0210] Among them, DC(i) represents the value of the degree centrality of node i;

[0211] k i is the degree of node i;

[0212] n is the number of nodes;

[0213] K-shell: The K-shell algorithm measures the centrality and propagation ability of nodes in the network by decomposing the network into multiple shells by removing nodes with degrees less than or equal to k layer by layer.

[0214] VC: The VC index measures the propagation influence of a node by calculating the number of communities connected to each node. The more communities a node is connected to, the greater the propagation influence of that node. This invention patent uses the louvain community detection algorithm to detect communities.

[0215] Semi-local centrality SLC: Semi-local centrality SLC is a method for evaluating the importance of nodes that combines local and neighborhood information, and is calculated based on the degree of a node and the degrees of its direct neighbors. Its calculation method is given by Equation (1).

[0216] RCNN: The RCNN algorithm constructs a feature matrix as the input of the convolutional neural network through the BFS algorithm, and calculates the importance score of the nodes through two layers of convolutional neural networks.

[0217] M-RCNN: The M-RCNN algorithm constructs the input of a three-channel convolutional neural network and obtains the importance score of the nodes through two convolutional layers, two pooling layers and a fully connected layer network.

[0218] Graph ATtention network (GAT): Based on the GNN, GAT designs a multi-head attention mechanism for the aggregation of the GNN by adaptively assigning different weights to neighbors.

[0219] 3.3 Performance Metrics

[0220] To comprehensively evaluate the ranking quality and importance correlation, we use three metrics, Kendall's τ, Monotonicity Index (MI), and Overlap Rate (OVER-K), to reflect the effectiveness of the ranking task. The formal definitions of the metrics are as follows.

[0221] Kendall's τ measures the consistency of the rankings of two variables, based on the comparison of node pairs. It evaluates the ranking consistency by calculating the number of node pairs that are in agreement and disagreement. Kendall's τ ranges from -1 to 1. The larger the τ value, the more consistent the two rankings are; the smaller the τ value, the less consistent the rankings are. The τ value is defined as:

[0222]

[0223] where τ represents the ranking measurement metric;

[0224] N L is the number of nodes in the sorted list, N c and N d are the numbers of node agreement pairs and disagreement pairs, respectively. For example, let (x 1 , y 1 ), (x 2 , y 2 ),..., (x n , y n ) be a set of joint ranks of X and Y, where X and Y represent two sets of ranking lists. If x i > x j and y i > y j or x i < x j and y i < y j , then (x i , y i ), and (x j , y j ) is a consistent pair. If x i > x j and y i < y j or x i < x j and y i > y j , then this pair is called an inconsistent pair. If x i = x j or y i = y j , then this pair is neither consistent nor inconsistent. Subsequently, the τ value will be used to quantify the correlation between the ranking list obtained by the SIR model and the ranking lists obtained by the baseline and MSACNN.

[0225] Monotonicity Index (MI). Since some nodes may have the same ranking, MI can be used to quantify the distinguishability of nodes in the entire ranking list. The range of MI is from 0 to 1. The larger the MI value, the higher the distinguishability of nodes in the sorting. The calculation formula of MI is:

[0226]

[0227] where MI represents the monotonicity index of the ranking result;

[0228] N L is the number of nodes in the ranking list, L 0 is the list of ranking values, α ∈ L 0 is the ranking value, N α is the number of nodes with the same ranking value α. The value range of MI is from 0 to 1. The larger the MI value, the higher the distinguishability of nodes in the ranking list. If MI = 1, then each node has a ranking value. If MI = 0, then all nodes have the same ranking value

[0229] OVER-K is an index that measures the overlap degree between the key nodes selected by an algorithm and the actually important nodes. Specifically, the coverage rate is used to evaluate what proportion of the top k selected nodes are actually key nodes. The calculation formula of OVER-K is:

[0230]

[0231] where OVER-K represents the index that measures the overlap degree between the key nodes selected by an algorithm and the actually important nodes;

[0232] is the set of the top k importance prediction nodes;

[0233] ∩ represents the intersection;

[0234] S top-k is the set of the top k true important nodes;

[0235] represents the intersection of the set of the top k predicted important nodes and the set of the top k true important nodes;

[0236] 3.4 Baseline Experiments

[0237] 3.4.1 Node Ranking Ability

[0238] To comprehensively evaluate the performance of MSACNN on different datasets, we show the Kendall’s τ, MI, and OVER-K values of MSACNN and other baseline methods on nine real datasets, as shown in Table 3. Since the proposed MSACNN method integrates information from three dimensions: node local, community, and global, the results are better than ordinary network structure-based centrality methods. Secondly, the multi-scale convolutional network with channel attention can well integrate multi-dimensional information. The hidden representation of each node obtained by MSACNN integrates many features, which improves the way of obtaining features by M-RCNN. By aggregating the neighbor degrees, the problem of feature acquisition on unconnected sparse graphs is solved. In addition, the weights of the multi-scale layers are trained by the Adam optimizer, breaking through the difficulty of manually assigning weights.

[0239] Table 3 Node Ranking Ability. Experimental results of eight different methods on nine real networks. Here, the neighborhood size S = 24 and the infection probability μ = 1.2μ c , and the experimental results with K = 100. The highest result is shown in bold, and the second highest result is underlined.

[0240]

[0241]

[0242] The final results show excellent performance in all three metrics. The improvement rates of the Kendall's τ value compared to the sub-optimal algorithms on each baseline are as follows: 2.2% on the Arenas network, 1.0% on the GrQ network, 3.2% on the Pgp network, 3.7% on the Hep network, 0.2% on the Stelzl network, 2.7% on the Figeys network, 10.07% on the PowerGrid network, and 4% on the Ns network. In terms of the monotonicity index, the centrality method based on network structure is likely to give the same score to different nodes, so the monotonicity index is not high, especially for the KS algorithm, because it decomposes the network into multiple shells and nodes in each shell have the same importance score. The deep learning-based method, on the other hand, shows excellent performance in the monotonicity index. Except for the Hep network, the monotonicity index of MSACNN is higher than or equal to that of other baseline methods. In terms of the coverage OVER-K, MSACNN also has excellent performance. This is because the fusion of multi-scale convolution and attention mechanism can extract hidden node information, enabling MSACNN to better identify key nodes. Except for the GrQ network, MSACNN's performance in the other eight real networks is improved by 1% to 13% compared to other baseline methods.

[0243] 3.4.2 Kendall's Coefficient at Different μ

[0244] To evaluate the performance of MSACNN on different real networks, parametric analysis was conducted using different infection rates μ. Figure 2 shows the Kendall's τ coefficient between the ranking score and the true infection scale, where μ c = <k> / ( <k2> - <k>) The ranking accuracy of the proposed MSACNN algorithm and other baseline methods in 9 different real networks was compared by Kendall's τ coefficient as Figure 2 shown. Under different infection probabilities μ, the node ranking accuracy of the proposed method is higher than that of other baseline methods on the vast majority of networks.

[0245] Figure 2 is a schematic diagram of the sorting ability of nodes. The sorting scores obtained by simulating the SIR model with different infection rates μ of DC, KS, VC, SLC, RCNN, M-RCNN, GAT, and MSACNN on nine real networks were compared, where μ c = <k> / ( <k2> - <k>), RCNN, M-RCNN, and MSACNN were trained using S = 24.

[0246] Specifically, except on the jazz network, this is because the jazz network is small in scale and has a large average degree, and the semi-local centrality effect is excellent. MSACNN performs much better than other algorithms on the PowerGrid and NS networks because the sparsity of these two networks is relatively high, and it is difficult for ordinary methods to obtain sufficient node importance information. MSACNN fills in the missing values by using the average neighbor degree when constructing the feature matrix, so it performs well on unconnected and sparse networks. It can be inferred that MSACNN will maintain stability when selecting neighborhood networks of different sizes to obtain features, especially when the size of the neighborhood network increases.

[0247] 3.4.3 Sorting Ability of Different Neighborhood Sizes

[0248] To study the influence of the neighborhood network scale S on the performance of the M-RCNN algorithm and provide guidance for adjusting S when applying this algorithm to networks with different structural features, we conducted the following analysis. We simulated the sorting scores obtained from the SIR model with the infection rate set to μ = 1.2μ c and the predicted sorting scores of the MSACNN algorithm trained with different neighborhood network scales S (from 8 to 64) on four disconnected and sparse networks, and evaluated the consistency between the two by calculating the Kendall's τ coefficient.

[0249] Figure 3 is a schematic diagram showing the influence of the neighborhood network size on the algorithm. It shows the comparison of the sorting scores obtained by RCNN, M-RCNN, and MSACNN on four disconnected or sparse networks with the SIR model when training the network with different sizes of S. Here, the infection probability is set to μ = 1.2μ c .

[0250] Figure 3 The results of c show that on these networks, the performance of RCNN and M-RCNN will decrease as the neighborhood network size S increases, especially on unconnected networks such as the PowerGrid network and the NS network, which is caused by their feature extraction methods. The proposed MSACNN still maintains good performance under different neighborhood network sizes S because it uses the method of aggregating neighbor degrees to replace zero-padding. On the GrQc and Hep networks, the performance of RCNN and M-RCNN also decreases as S increases. The stability of the proposed algorithm can be proven through the above results.

[0251] 3.4.4 Ablation Experiment

[0252] To evaluate the functions of each component in the MSACNN algorithm, we divide the entire algorithm into the following components and combine them pairwise to form three variants: the component for constructing three-channel input features, the multi-scale convolution component, and the channel attention component. Under the condition of S = 24, the Kendall's τ coefficient with the sorting results of the SIR model at different infection probabilities μ is calculated and compared with the performance of the complete MSACNN algorithm.

[0253] · AT+MS: Only use the channel attention mechanism and multi-scale convolution. Construct a single-channel input and use degree values to convert the adjacency matrix into a feature matrix.

[0254] · TH+AT: Only use three channels and the channel attention mechanism. The convolutional layer uses a convolutional kernel of size 5*5.

[0255] · TH+MS: Only use three channels and multi-scale convolution. Remove the channel attention mechanism and construct a local, community, and global three-channel input multi-scale convolutional neural network.

[0256] Figure 4 It is a schematic diagram of the experimental results of using different component combinations in the algorithm. AT+MS: Only use the channel attention mechanism and multi-scale convolution. TH+AT: Only use three channels and the channel attention mechanism. TH+MS: Only use three channels and multi-scale convolution. The abscissa is different infection probabilities μ, where μ c = <k> / ( <k2> - <k>), where different variants are all trained with S = 24.

[0257] From Figure 4 It can be seen that on connected or unconnected real networks, MSACNN is always superior to other variants, verifying the combined advantages of the components used. Thus, it can be seen that constructing a three-channel multi-scale convolutional attention network has advantages in identifying key nodes. Variants considering only two components are still unstable when dealing with different types of networks, especially on unconnected networks. TH+AT and TH+MS are basically superior to AT+MS, indicating that the feature extraction method we constructed, which integrates node local, community, and global information, is very effective. Through these adjustments, the MSACNN algorithm can be further optimized in future applications to better adapt to various complex network environments.

[0258] 3.4.5 Correlation Analysis

[0259] To evaluate the discriminative ability of the proposed algorithm and the M-RCNN algorithm, we show the correlation of the node ranking scores obtained by the MSACNN algorithm and the M-RCNN algorithm. Specifically, the MSACNN algorithm and the M-RCNN algorithm uniformly use S = 24 in all networks. Figure 5 In, the horizontal axis and the vertical axis respectively represent the normalized diffusion influence of nodes calculated by the MSRCNN and M-RCNN algorithms, and the color of each point represents the node diffusion influence obtained by simulating the SIR model with μ = 1.2μ c at this time.

[0260] Figure 5 is the correlation of the normalized propagation influence scores predicted by the MSACNN algorithm and the M-RCNN algorithm on 9 real networks. The color of each point is determined by the infectious influence calculated by simulating the SIR model, and the infection rate is μ = 1.2μ c , S = 24. The points on the diagonal indicate that the two algorithms have the same propagation influence score for the same node; the points above the diagonal indicate that the MSACNN algorithm has a higher propagation influence score for this node than the M-RCNN algorithm, while the points below the diagonal indicate that the M-RCNN algorithm has a higher score. By observing the color and position of the points, the abilities of the two algorithms in identifying propagation influence nodes can be compared.

[0261] In Figure 5 each subgraph of, a diagonal line is also added to more clearly show the discriminative ability of the two algorithms. The points above the diagonal indicate that the MSACNN algorithm assigns a higher diffusion influence score to the corresponding node than the M-RCNN algorithm, and vice versa. By analyzing the distribution of points and color changes on both sides of the diagonal, we can compare the differences between these two algorithms in node discrimination. As Figure 5 As shown, in the 9 real networks, the ranking scores of the MSACNN algorithm and the M-RCNN algorithm are both positively correlated. This indicates that both algorithms have good capabilities in discriminating node importance. At the same time, it can be seen from the figure that except for the stelzl network, the MSACNN algorithm can better identify nodes with greater influence than the M-RCNN algorithm, indicating that the proposed algorithm is superior to the M-RCNN algorithm in the discriminative ability of key nodes. This is because the MSACNN algorithm improves the deficiencies in feature extraction compared to the M-RCNN algorithm. For example, it uses the K-shell algorithm as the feature extraction method and addresses the feature problem of unprocessed disconnected networks. At the same time, it also incorporates a multi-scale convolutional neural network and a channel attention mechanism to better mine node importance information.

[0262] 4 Conclusion

[0263] How to identify key nodes in complex networks has always been an important issue in many fields. Based on deep learning methods and considering local, community, and global structure information simultaneously, this invention patent proposes an MSACNN algorithm based on GCN, which solves the problems of the single-sidedness of previous centrality algorithms based on network structure and the difficulty of constructing features based on deep learning algorithms. The proposed algorithm embeds local, community, and global information into the feature matrix and uses the method of aggregating neighbor degrees to supplement information for nodes with insufficient neighbors. The constructed three-channel information can capture the hidden information of nodes under the combined action of multi-scale convolutional kernels and channel attention, thus achieving an efficient node importance recognition task. Experimental results show that on nine different real networks, MSACNN performs better than seven other baseline methods. In terms of information propagation, MSACNN improves by 0.2% to 10.07% compared to other baseline methods, and at the same time has better discriminability and key node prediction accuracy than baseline methods. Experiments under different neighborhood sizes show that MSACNN is more stable than other methods, and the effectiveness of each component in the algorithm is verified through ablation experiments. In addition, MSACNN has a low time complexity and can be used for large-scale networks.

[0264] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.< / k> < / k2> < / k> < / k> < / k2> < / k> < / k> < / k2> < / k> < / k> < / k> < / k> < / k> < / k> < / k> < / k>

Claims

1. A method for identifying key nodes in complex networks based on graph convolutional neural networks, characterized in that: The following steps are involved: S1, obtain the neighborhood network of each node; S2, obtain local, community, and global information based on the network; S3, inputs data information into the multi-scale network; S4, outputs the identified key nodes.

2. The method for identifying key nodes of complex networks based on graph convolutional neural networks according to claim 1, characterized in that: The method for generating local information in step S2 is: Where SLC(i) represents the value of the semi-local centrality of node i; Γ(i) is the set of direct neighbors of node i; Γ(j) is the set of direct neighbors of node j; Q(j) represents the sum of the degrees of the second-order neighbors of node j; N(w) is the number of second-order neighbors of node w; Among them, X Loc Represents the local hierarchical information matrix of the node; A ij Represents the adjacency matrix of the network; SLC(j) represents the value of the semilocal centrality of node j; r represents the number of neighbors of the node; md is the neighbor aggregation degree of the corresponding node; S represents the neighborhood size; SLC(i) represents the value of the semilocal centrality of node i.

3. The complex network key node identification method based on graph convolutional neural network according to claim 1 is characterized in that: The method for generating community information in step S2 is: VC(i)=|V(Γ(i))|, Among them, VC(i) represents the importance value of node i in the community; V(Γ(i)) represents the set of communities to which node i and its neighbors belong; |V(Γ(i)| represents the number of community sets to which node i and its neighbors belong; Among them, X Com Represents the community level information matrix of nodes; A ij Represents the adjacency matrix of the network; VC(j) represents the importance value of node j in the community; r represents the number of neighbors of the node; md is the neighbor aggregation degree of the corresponding node; S represents the neighborhood size; VC(i) represents the value of the importance of node i in the community.

4. The method for identifying key nodes of complex networks based on graph convolutional neural networks according to claim 1, characterized in that: The method for generating global information in step S2 is: Among them, EC(i) represents the value of the eigenvector centrality of node i; x i is the eigenvector centrality value of node i; c is the proportionality constant; n represents the number of nodes; a ij Represents the element in the i-th row and j-th column of the adjacency matrix; x j represents the eigenvector centrality value of node j; Among them, X Glo Represents the full hierarchical information matrix of the node; A ij Represents the adjacency matrix of the network; EC(j) represents the value of the eigenvector centrality of node j; r represents the number of neighbors of the node; md is the neighbor aggregation degree of the corresponding node; S represents the neighborhood size; EC(i) represents the value of the eigenvector centrality of node i.

5. The method for identifying key nodes of complex networks based on graph convolutional neural networks according to claim 1, characterized in that: Step S3 includes: Among them, DC(i) represents the value of degree centrality of node i; k i is the degree of node i; n is the number of nodes.

6. The complex network key node identification method based on graph convolutional neural network according to claim 1 is characterized in that: The multi-scale network consists of three convolutional layers, three pooling layers and one fully connected layer.

7. The method for identifying key nodes of complex networks based on graph convolutional neural networks according to claim 6, characterized in that: After step S4, the method further includes step S5, which verifies the indicators of the multi-scale network.

8. The complex network key node identification method based on graph convolutional neural network according to claim 6 is characterized in that: The indicators include one of the ranking measure, the monotonicity index of the ranking results, and the overlap degree index, or any combination thereof.

9. A computer system, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the complex network key node identification method based on graph convolutional neural network as described in one of claims 1 to 8 when executing the executable instructions.

10. A computer-readable storage medium, characterized in that: include: a memory having a computer program stored thereon; A processor is used to execute the program in the memory to implement the complex network key node identification method based on graph convolutional neural network as described in any one of claims 1 to 8.