Key node identification method and system of traffic transportation network
By destroying the local network structure, using KL divergence to analyze the changes in the information transmission probability distribution before and after node removal, and combining global and local attributes to evaluate the destructive impact of nodes, the problem of difficult determination of node importance in transportation networks is solved, and accurate identification and decision support of key nodes are achieved.
Patent Information
- Application Number
- CN202510973726.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-23
AI Technical Summary
When nodes in existing transportation networks are damaged by environmental factors, it is difficult to accurately determine the importance of nodes in the network, and traditional algorithms are not applicable enough in real networks.
By destroying the local network structure, the KL divergence is used to analyze the changes in the information transmission probability distribution before and after the node is removed. Combining the global and local properties of the node, the destructive impact of the node is evaluated and the key nodes are identified.
It improves the accuracy and applicability of key node identification, is applicable to large-scale transportation networks, and provides effective route planning and traffic flow allocation decision support.
Smart Images

Figure CN120692177A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network key node identification, and in particular relates to a method and system for identifying key nodes in a transportation network. Background Art
[0002] In transportation networks, the network environment is volatile and vulnerable to disruption. Damage to the network structure can lead to changes in transmission efficiency. Therefore, identifying key nodes in the network can effectively support decision-making for route planning, traffic flow allocation, and site construction. Key node identification methods have been widely applied in various complex networks, providing significant research value and application significance for analyzing the security, survivability, and stability of network systems.
[0003] A large number of algorithms have emerged for identifying key nodes. The main classic algorithms include degree centrality (DC), betweenness centrality (BC), closeness centrality (CC), eigenvector centrality (EC), and Kshell. However, each of these algorithms has its advantages and disadvantages. The DC algorithm considers relatively few factors; the BC and CC algorithms, which rely on shortest path statistics, have high time complexity; the EC algorithm has certain limitations in large-scale and heterogeneous networks; and the Kshell algorithm uses a layer-peeling method to calculate the position of a node in the entire network, but its node differentiation is coarse-grained. Therefore, there is still a need for a key node identification method that is simple to implement, highly accurate, and widely applicable.
[0004] However, in the process of information transmission in transportation networks, information is transmitted between nodes through connections. When a node in the network is damaged and fails, the information transmission path between nodes in the local network will change. Moreover, when different nodes fail, the destructive impact on the network is different. Most of the algorithms currently used analyze the multiple attributes of nodes while keeping the network structure unchanged. However, in real networks, nodes are often damaged by environmental factors, making it difficult to accurately determine the importance of nodes in transportation networks. Summary of the Invention
[0005] In response to the deficiency in existing transportation networks that it is difficult to accurately determine the importance of nodes in the network when nodes are damaged by environmental factors, the present invention provides a method and system for identifying key nodes in transportation networks. By destroying the local network structure, the node and its one-hop neighbor nodes and connecting edges are removed. From the perspective of information transmission, the KL divergence is used to analyze the impact of the node on the local network before and after removal. At the same time, the global and local attributes of the node are integrated to measure the destructive impact of the node, thereby achieving an assessment of the node's importance, thereby solving the problems existing in the existing technology.
[0006] A method for identifying key nodes in a transportation network comprises the following steps: Obtain all nodes in the transportation network structure and their Kshell values and degree values, and select one of the nodes; Based on the degree values of all nodes, the true probability distribution of information transmission between the selected node’s neighboring nodes and the selected node is calculated, as well as the fitted probability distribution of information transmission between the selected node and the one-hop neighboring nodes in the neighborhood after removing the edges connecting the selected node and the surrounding neighboring nodes. Based on the true probability distribution and the fitted probability distribution, the KL divergence model is used to obtain the information entropy lost by each neighboring node of the selected node. Based on the information entropy lost by the neighboring node and the Kshell value, the KL divergence value - the information entropy lost after the selected node is removed for all one-hop neighboring nodes in the neighborhood is obtained. Based on the KL divergence value - the information loss entropy and the degree value of the selected node itself, the destructive impact of the selected node is obtained. Based on the destructive impact of the selected node and its one-hop neighboring nodes in the neighborhood, the importance value of the selected node is obtained. By calculating the importance value of each node, the key nodes of transportation network structure damage are identified.
[0007] Furthermore, the true probability distribution of neighboring nodes of the selected node transmitting information through the selected node is expressed as: ; in, n Represents the node number in the network, p i ( j ) represents the selected node i No. j The probability of neighbor nodes transmitting information, j =1,2,3... n is a node i Neighbor nodes; information transmission probability p i ( j ) is equal to the neighbor node j The degree value and node i The proportion of the sum of the degree values of all neighbor nodes, which is expressed as: ; in, d(j) Representative neighbor node j The degree value, Γ ( i ) represents a node i The set of all one-hop neighbor nodes.
[0008] Furthermore, after removing the edges connecting the selected node and the surrounding neighboring nodes, the fitted probability distribution of information transmission between one-hop neighboring nodes in the neighborhood is expressed as: ; in, in Representative Node i No. n neighbor nodes, q in ( x ) represents neighbor nodes n Respectively with nodes i The fitting probability distribution of information transmission among all one-hop neighbor nodes; q in ( x ) is calculated as: ; in: ; in, Representative Node i After removal, its neighbor nodes n With neighboring nodes j Fitting probability distribution for information transmission, Calculated based on the shortest distance and degree between nodes, Representative Node n and j The influence coefficient of the information transmission distance between Representative Node i After the destruction, its neighbor nodes j The current degree value.
[0009] Furthermore, the KL divergence model is used to obtain the information entropy D' lost by each neighbor node of the selected node based on the true probability distribution and the fitted probability distribution. ij The formula for calculating (P||Q) is: .
[0010] Furthermore, according to the information entropy and Kshell value of the neighbor node loss, the KL divergence value generated by the selected node after removal for all one-hop neighbor nodes in the neighborhood - the information entropy loss D i The formula for calculating (P||Q) is: .
[0011] Furthermore, the destructive influence of the selected node is obtained according to the KL divergence value-information loss entropy and the degree value of the selected node itself. DE ( i), which is calculated as follows: ; in, d ( i ) represents a node i The degree value.
[0012] Furthermore, the importance value of the selected node is obtained based on the destructive influence of the selected node and the one-hop neighboring node in the neighborhood. KLN ( i ), which is calculated as follows: .
[0013] The present invention also proposes a key node identification system for a transportation network, comprising: The acquisition module is used to obtain all nodes in the transportation network structure and their Kshell values and degree values, and select one of the nodes; The probability distribution calculation module is used to calculate the true probability distribution of information transmission between the neighboring nodes of the selected node and the selected node based on the degree values of all nodes, and the fitted probability distribution of information transmission between the one-hop neighboring nodes in the neighborhood after removing the edges connecting the selected node and the surrounding neighboring nodes; Importance value calculation module, used to obtain the information entropy lost by each neighboring node of the selected node using the KL divergence model based on the true probability distribution and the fitted probability distribution; obtain the KL divergence value minus the information entropy lost by each neighboring node in the neighborhood after the selected node is removed based on the information entropy lost by the neighboring node and the Kshell value; obtain the destructive impact of the selected node based on the KL divergence value minus the information entropy lost and the degree value of the selected node itself; and obtain the importance value of the selected node based on the destructive impact of the selected node and its one-hop neighboring nodes in the neighborhood; The key node identification module is used to identify the key nodes of transportation network structure damage by calculating the importance value of each node.
[0014] The present invention provides a method for identifying key nodes in a transportation network, which has the following beneficial effects: The present invention removes nodes from the transportation network and uses the damage caused by the nodes to the local network structure. From the perspective of information transmission between nodes, the present invention analyzes the changes in the probability distribution of communication between nodes in the local network structure before and after the node removal, and uses the idea of KL divergence to measure the information entropy lost before and after the local network change. The destructive impact of the node removal on the network is evaluated by combining the node's own global network attributes and its own local attributes. Finally, by calculating the destructive impact of all nodes in the local network structure centered on the node, the final determination of the node's importance is achieved, and the accuracy and effectiveness of the identification results are improved. The present invention combines the characteristics of the actual transportation network environment, such as being volatile and easily damaged, and analyzes and identifies key nodes in the network by simulating the destruction of the network structure. It can be well applied to transportation networks of most sizes and other types of complex network structures, improves the accuracy and effectiveness of important nodes, and provides effective decision support for route planning, traffic flow distribution, site construction, etc. of the transportation network. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a structural diagram of the network model in the present invention; Figure 2 It is a schematic flow chart of the algorithm of the present invention; Figure 3 Schematic diagram of the process of solving KL divergence in the algorithm of the present invention; Figure 4 This is an example diagram of the importance of nodes in the present invention; Figure 5 is a comparison chart of Kendall's ɩ results for Karate, Dolphins, Football, Euroroad, Usair, and Friendships in an embodiment of the present invention; Figure 6 This is a comparison chart of the Kendall ɩ results of Protein, Powergrid, HepPh, and Ca-astroph in the examples of the present invention; Figure 7 is a comparison chart of node infection capability curves of the algorithm of the present invention and Karate, Dolphins, Football, and Euroroad in an embodiment of the present invention; Figure 8 is a comparison chart of node infection capability curves of the algorithm of the present invention and Usair, Friendships, Protein, and Powergrid in an embodiment of the present invention; Figure 9 is a comparison diagram of the node infection capability curves of the algorithm of the present invention and HepPh and Ca-astroph in an embodiment of the present invention; Figure 10This is a comparison chart of the infectious capacity curves of the top ten important nodes of Karate, Dolphins, Football, and Euroroad in an embodiment of the present invention; Figure 11 This is a comparison chart of the infectious ability curves of the first ten important nodes, namely, Usair, Friendships, Protein, Powergrid, HepPh, and Ca-astroph, in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0017] The present invention proposes a key node identification method for transportation network (KLN), such as Figure 1 As shown, by destroying the local network structure, the node, its one-hop neighbor nodes, and connecting edges are removed. From the perspective of information transmission, the KL divergence (relative entropy) is used to analyze the impact of the node removal on the local network before and after the node is removed. At the same time, the global and local attributes of the node are integrated to measure the destructive impact of the node, thereby evaluating the node's importance. Key node identification technology, a hot research topic in complex networks, has been applied to urban transportation networks in recent years. Using key node identification methods can help study and analyze the structure of real networks. By identifying key hubs in urban transportation networks, it can effectively conduct network robustness and anti-destruction research, and avoid cascading accidents caused by the destruction of certain key nodes. Therefore, key node identification methods are of great significance in the study of urban transportation networks and can bring good economic benefits and application value.
[0018] In the research on key node identification methods, how to design accurate, fast, and efficient methods to identify key nodes in complex networks has always been a hot topic pursued by scholars. A large number of algorithms have emerged, including classic algorithms such as degree centrality (DC), betweenness centrality (BC), closeness centrality (CC), eigenvector centrality (EC), and Kshell. However, each of these algorithms has its advantages and disadvantages. The DC algorithm considers relatively few factors; the BC and CC algorithms, which rely on shortest path statistics, have high time complexity; the EC algorithm has certain limitations in large-scale and heterogeneous networks; and the Kshell algorithm uses a layer-peeling method to calculate the position of nodes in the entire network, but the node distinction is coarse-grained. Therefore, there is still a need for a key node identification method that is simple to implement, highly accurate, and widely applicable.
[0019] In recent years, a large number of new algorithms have been proposed based on the improvement and innovation of the above-mentioned classic algorithms. Most of them integrate multiple attributes of nodes, such as the degree value, Kshell value, clustering coefficient and path information of nodes, and integrate global attributes and local attributes to achieve the improvement of the effectiveness and accuracy of key node identification results, such as PL, GIN, ECRM and other algorithms. At the same time, based on the fusion of multiple attributes, a large number of algorithms have been deeply innovated. For example, the multi-attribute fusion based on the idea of gravity model can realize the identification of key nodes by introducing the attributes of nodes (degree value, Kshell value and the shortest distance between nodes) into the gravity model, such as Algorithms such as GSM, MCGM, and KSGC generally have high time complexity and are therefore unsuitable for large-scale networks. Multi-attribute fusion based on eigenvector centrality utilizes the relationship between nodes and their neighbors for comprehensive judgment, such as PR, Hits, and GSI algorithms. However, the applicability of these algorithms depends on the network structure. Multi-attribute fusion based on entropy centrality analyzes the relationship between nodes and neighbors in the network from the perspective of information transmission, converting the relationship between nodes and neighbors into information entropy analysis, such as Enrenew, MCDE, and LFIC. Multi-attribute fusion based on voting, such as VoteRank and VoteRank Plus, is primarily suitable for selecting top-N key nodes. In summary, how to innovate methods based on the integration of global and local attributes to design algorithms with high accuracy, simple execution, and strong applicability has always been a hot topic of concern in this field.
[0020] Through the research and analysis of the algorithms proposed in recent years, it is found that most of the algorithms are based on the analysis of multiple attributes of nodes under the condition that the network structure remains unchanged. In real networks, especially in transportation networks, transportation hub nodes are usually damaged by environmental factors. Therefore, the analysis idea of robustness in network structure is used to simulate the impact of node failure in real network systems on the network, so as to determine the importance of nodes in the network. In the process of information transmission in the network, information is transmitted between nodes through connections. When a node in the network is damaged and fails, the information transmission path between nodes in the local network will change, and when different nodes fail, the destructive impact on the network is different. Using Figure 1Take the toy transportation network model in [1] as an example, where nodes 2 and 5 are located in different Kshell layers. When node 2 is removed, nodes 1 and 3 have no path to reach node 6. When node 5 is removed, nodes 9 and 10 have no path to reach node 1. The entire toy transportation network is split into three disconnected sub-networks. From the perspective of information transmission, the removal of a node causes some of its neighboring nodes to lose some information due to the inability to transmit information, which has a certain destructive impact on the network. This impact is related to the basic properties of the network in which the node is located and the impact of the node failure on the information transmission of neighboring nodes. The greater the impact of node removal on the information transmission between neighboring nodes, the more destructive the node is. Therefore, how to use the destructive impact of node removal on the local network before and after it is removed to measure the importance of a node is a novel and very meaningful challenge.
[0021] In this invention, the definitions of degree centrality, K-shell, and KL divergence are applied. Assume a network G =( V , E ), V and E represent the number of nodes and edges in the network respectively.
[0022] (1) Degree centrality: Degree centrality is used to measure the number of neighbor nodes of a node and is the most direct way to measure the influence of a node. i The degree centrality of is calculated as:
[0023] (1) in, d ( i ) represents a node i The degree value of the node i The number of directly connected neighbor nodes, N represents the total number of nodes in the network.
[0024] (2) K-shell decomposition: The K-shell decomposition method measures the importance of a node's position in the network. By removing the outer nodes layer by layer, the inner nodes have a higher influence. When performing the K-shell decomposition, first remove the nodes with a degree of 1 and the edges connected to them. When new nodes with a degree of 1 appear in the remaining network, these nodes are removed again until there are no more nodes with a degree of 1 in the remaining network. All the removed nodes form a layer with a K-shell value of 1. Then, start removing nodes with a degree of 2, and repeat this process.
[0025] (3) Shannon entropy: Shannon information entropy is used to express the uncertainty of the system. The calculation formula is: (2) in N represents the number of possible values of a random variable; x represents a random variable; p ( x i ) represents a random variable x i When identifying key nodes, the Shannon information entropy of the node can be used to indicate its importance in the network. The larger the entropy value, the more disordered the state around the node, and the greater the uncertainty in information transmission between the node and other nodes. Increased uncertainty leads to an increase in the amount of information, so the larger the entropy value of the node, the more important the node.
[0026] (4) KL divergence: KL divergence is a measure of the asymmetry of the difference between two probability distributions, one of which is the true distribution. P ( X ), and the other is the theoretical (fitted) distribution Q ( X ), then the KL divergence (relative entropy) at this time is equal to the difference between the cross entropy and the true distribution information entropy, indicating the information loss generated when the theoretical distribution is used to fit the true distribution. KL divergence is calculated as:
[0027] (3) in, P ( x i ) is the probability distribution of real events P ( X ) x i The probability of an element, Q ( x i ) is the probability distribution of the event obtained by theoretical fitting. x i The larger the KL divergence value is, the more information entropy is lost in the probability distribution after fitting.
[0028] The method specifically comprises the following steps: S1. Get all nodes in the network structure and their Kshell Value and degree value. Assume an undirected and unweighted network , N represents the number of nodes in the network, E represents the number of edges, and the neighbor set Γ(v) of node v.
[0029] S2. Calculate the shortest distance. Define nodes i and j The shortest distance betweenl ( i , j ), calculated as:
[0030] (4) in, shortpathbetween ( i , j ) indicates the nodes in the entire network i and nodes j The number of hops of the shortest path among all paths.
[0031] Calculate the path influence coefficient. Define the node i and j The shortest distance between l ( i , j ), calculated as:
[0032] (4) in, l ( i , j ) indicates the nodes in the entire network i and nodes j The number of hops of the shortest path among all paths.
[0033] In the above formula, when a node in the network is removed, if its neighbor nodes i and j When there is still a path between the two nodes, the distance influence coefficient is equal to the inverse of the shortest distance between the two nodes. In addition, the number of hops for data communication between the node and itself is set to 1. i and j When there is no path for transmitting information, the distance influence coefficient is set to 1 / 2, and only the shortest distance between the nodes before they are removed is considered to be 2.
[0034] S3. Calculate the true probability distribution P(X). i Before the destruction, the node i The neighbor nodes of i Transmit information with other neighbors. The probability of neighbor nodes transmitting information is related to the proportion of their degree value in the sum of the degree values of all neighbor nodes of node i. The larger the value, the greater the probability of the neighbor node transmitting information with other neighbor nodes. Therefore, the node i The actual probability distribution of information transmission to the neighbors around is: P i ( x ) is defined as:
[0035] (5) in, n Represents the node number in the network, p i ( n ) represents a node i No. n The probability of information transmission between neighbor nodes is {1,2,3... n} is a node i neighbor nodes. p i ( n ) is equal to the neighbor node n The proportion of the degree value of the node to the sum of the degree values of all neighboring nodes is calculated by formula (7), where d ( n ) represents neighbor nodes n The degree value.
[0036] (6) S4. Calculate the fitting probability distribution Q(X). i After its connection edge is destroyed, the impact on each neighbor node is different. The probability of each neighbor node transmitting information with other neighbor nodes will change with the destruction of the network structure. Define the fitting probability distribution sequence of information transmission between all neighbor nodes of node i Q i ( x )for:
[0037] (7) in, in Representative Node i No. n neighbor nodes, q in ( x ) represents neighbor nodes n Respectively with nodes i All one-hop neighbor nodes (1,2,3... n ) to carry out the fitted probability distribution of information transmission, which is calculated as: (8) in: (9) Representative Node i After removal, its neighbor nodes n With neighboring nodes j Fitting probability distribution for information transmission, The calculation of is related to the shortest distance and degree between nodes. Representative Noden and j The influence coefficient of the information transmission distance between them is calculated by formula (5): d’ ( j ) represents a node i After the destruction, its neighbor nodes j The current degree value.
[0038] S5. Calculate KL divergence-information loss entropy. Using node i The true probability distribution P of information transmission between neighboring nodes before destruction i ( x ), and the pseudo-probability distribution of information transmission between neighboring nodes after destruction q ij ( x ), according to formula (11), the information entropy of each neighbor node loss is calculated , comprehensive neighbor nodes j Global properties of Kshell Value, get node i KL divergence of the loss of all neighboring nodes after destruction - information entropy D i (P||Q) is obtained from formula (12):
[0039] (10) (11) S6. Calculate the destructive impact of nodes. i The destructive effect of the node is caused by the KL divergence-information loss entropy on the surrounding neighbors after the node is destroyed, which is related to the node's own degree value. d ( i ) represents a node i The degree value of . Calculated as:
[0040] (12) S7, Calculation of node importance. i Importance by node i The destructive impact of the node itself and all its one-hop neighbor nodes is determined by the node i As the center, calculate the destructive impact value of all one-hop neighbor nodes in the local network structure, and finally get the importance value of the node KLN ( i ), calculated as shown in formula (14):
[0041] (13) After the network is initialized, first select a node to be removed i,When executing the algorithm, it is mainly divided into 6 steps, Step 1: Count the global and local attributes of the node; Step 2: Calculate the true probability distribution P(X) of the information transmission of the neighboring nodes around node i; Step 3: After removing node i, calculate the fitting probability distribution Q(X); Step 4: Calculate the node i The KL divergence before and after removal; Step 5: Calculate the destructive influence of the node. After the calculation is completed, traverse the other nodes in the network in turn until the destructive influence values of all nodes are obtained. Step 6: Calculate the importance ranking value of the node. The execution of the above steps is as follows Figure 2 shown.
[0042] Below is Figure 3 Taking the toy transportation network as an example, node 1 in the network is selected as the removed node, and each step in the execution process of this method is illustrated with examples.
[0043] Step 1: Calculate the local and global properties of the node. Calculate the degree value of node i and the K-shell value of node i, such as Figure 3 As shown in (a), the degree of node 1 is 4. Kshell The value is 2, and the neighbor nodes of node 1 are: 2, 6, 7, and 8.
[0044] Step 2: Compute Node i of P ( x ). According to the information transmission connection relationship between nodes, Figure 3 As shown in (b), the four neighboring nodes of node 1 can all transmit information through node 1. According to formulas (6) and (7), the probability distribution of information transmission between neighboring nodes of node i can be calculated as:
[0045] (14) Step 3: Compute Node i of Q ( x ). Remove node 1 and its connected edges. Take node 7, the neighbor of node 1, as an example. Figure 3 As shown in (c), when node 7 transmits information with its neighboring nodes 2, 6, 7, and 8, the shortest distance between node 7 and other nodes and the influence coefficient of the distance are first obtained according to formulas (4) and (5). Among them, the shortest distances between node 7 and nodes 2, 6, 7, and 8 are 4, 3, 1, and 2 respectively. Therefore, using formulas (8), (9), and (10), we can calculate:
[0046] (15) (16) Calculate the information transmission probability of all neighbor nodes of node 1 and its neighbors in sequence, and finally obtain The distribution is: (17) Step 4: Calculate the KL divergence value of the node. Figure 3 As shown in (d), according to formula (11), P1( x ) and Q1( x ) is the fitted distribution of each neighbor node in q in ( x ) Calculate the information entropy loss caused by the removal of node 1 , according to the neighbor nodes j Global properties in the network Kshell The final KL divergence value of node 1 is:
[0047] (18) Step 5: Calculate the destructive impact of the node. In the toy transportation network, the degree of node 1 is equal to 4. According to formula (13), the final destructive impact of node 1 is DE(1) = 0.218 + 4 = 4.218.
[0048] Step 6: Calculate the importance value of the node. After the destructive value of all nodes in the entire network is calculated, take node 1 as the center and synthesize the destructive impact value of node 1 and its one-hop neighbors 2, 6, 7, and 8 in the local network structure, as shown in the following figure: Figure 4 shown.
[0049] According to formula (14), the final critical ranking value of node 1 is calculated to be: (19) In the present invention, the importance ranking results and values of all nodes are shown in Table 1.
[0050] Table 1 Importance ranking results and values of all nodes Experiments and Analysis: All experiments were run on a desktop computer running Windows 10, with an i3-10100 CPU and 8GB of RAM. Ten representative real-world network datasets were used to compare the proposed algorithm with several early classic algorithms, including DC, K-shell, EC, and PR, as well as more recently proposed algorithms such as PL, Enrenew, MCDE, ECRM, GSM, GSI, and MCGM. The performance of the proposed algorithm was verified through analysis of the experimental results. The algorithm was evaluated on ten representative real-world networks: the social networks Karate, Dolphins, and Friendships; the interaction network Football; the transportation and circuit transmission networks Euroroad, Usair, and Powergrid; the protein network Protein; the paper citation network HepPh; and the collaboration network Ca-astroPh. Table 2 shows the relevant characteristics of these real-world networks.
[0051] Table 2 Statistics of relevant characteristics of real networks We selected four networks (Dolphins, Euroroad, Protein, and Powergrid) from among ten networks with varying node counts. We compared the top ten nodes ranked by the 12 algorithms and analyzed their accuracy. As shown in Tables 3-6, when comparing the KLN algorithm, nodes in the 11 compared algorithms with the KLN algorithm in the same order are underlined.
[0052] Table 3 Top ten important nodes of 12 algorithms in Dolphins network Table 4 Top ten important nodes of 12 algorithms in Euroroad network Table 5 The top ten important nodes of the 12 algorithms in the protein network Table 6 Top ten important nodes of 12 algorithms in Powergrid network As can be seen in Table 3, the KLN algorithm and the DC algorithm all have the same 10 nodes, the MCDE algorithm has 9 nodes in common, the MCGM, GSI, ECRM, Enrenew, EC, PL, and PR have 8 nodes in common, and the GSM and Kshell algorithms have 5 and 3 nodes in common, respectively. In addition, the KLN algorithm has the highest similarity with MCGM in node order, with 5 nodes in the same order.
[0053] Table 4 shows that in the Euroroad network, KLN shares nine nodes with the DC, PL, Enrenew, MCDE, and GSI algorithms, eight with the ECRM algorithm, six with the EC and PR algorithms, five with the MCGM and GSM algorithms, and two with the Kshell algorithm. This is because the coarse-graining of the Kshell algorithm leads to a decrease in node differentiation as the number of network nodes increases. Furthermore, in terms of node order consistency, the KLN algorithm has the highest similarity with the GSI algorithm, with four nodes being in the same order.
[0054] As shown in Table 5, in the Powergrid network, KLN shares eight nodes with MCDE, Enrenew, and PL, six with ECRM, five with MCGM, GSI, EC, and DC, and four, three, and two with GSM, Kshell, and PR, respectively. Furthermore, in terms of node ordering, KLN has the highest similarity with GSI and MCDE, with two nodes sharing the same order.
[0055] As shown in Table 6, in the HepPh network, the KLN algorithm shares all 10 nodes with the MCGM and PL algorithms, 9 nodes with the GSI, EC, and ECRM algorithms, 8 nodes with the DC, MCDE, and GSM algorithms, 7 nodes with the Enrenew algorithm, 5 nodes with the PR algorithm, and 2 nodes with the Kshell algorithm. Furthermore, in terms of node ordering, the KLN algorithm has the highest similarity with the MCGM algorithm, with 8 nodes sharing the same order.
[0056] Kendall's ɩ model: Kendall's ɩ is used here to demonstrate the accuracy of the algorithm's results and the ranking results using the SIR model. The SIR model is a commonly used epidemic model that effectively simulates the entire process of infection transmission between nodes. In this model, nodes are divided into three categories: susceptible nodes (I), infectious nodes (S), and recovered nodes (R). During each time period, when an S node comes into contact with an I node, the I node becomes an S node with an infection probability of λ. Simultaneously, the S node becomes an R node with a recovery rate of μ, and the R node no longer participates in the infection process. Leveraging the contagion relationship in the SIR model, several nodes in the network are first set as S nodes with a certain initial probability. The number of infected nodes over a period of time is counted to represent the node's infectious capacity. In the experiments, the SIR infection probability λ was set between 0.01 and 0.1 to avoid excessive values of λ, which would lead to rapid infection and ineffective evaluation of individual node importance. The final result is the average value after 1000 iterations.
[0057] Based on the node importance ranking results obtained by the SIR model, the correlation analysis between the ranking results of the algorithm in the present invention and the comparison algorithm and the SIR results was performed. The Kendall ɩ value was used to reflect the accuracy of the algorithm results. The larger the value, the more similar the ranking results are to the results of the SIR model, and the higher the accuracy of the algorithm. The calculation formula of Kendall ɩ is as follows:
[0058] (20) In formula (20), ɩ( X , Y ) is used to determine the sequence X and sequence Y Similarity of elements. n is the total number of elements, n c and n d Represents two sequences respectively X and Y The number of consistent and inconsistent logarithms. Figure 5 Kendall comparison of the ranking results of the algorithm of the present invention and the comparison algorithm under 10 networks with the results of the SIR model .
[0059] Depend on Figure 5 、 Figure 6 As shown in Figure 1, a total of ten networks were selected, such as Figure 5 (a) Karate, (b) Dolphins, (c) Football, (d) Euroroad, (e) Usair, (f) Friendships, and Figure 6(g) Protein, (h) Powergrid, (i) HepPh, (j) Ca-astroph; KLN has the best overall effect in Dolphins, Euroroad, Protein, Powergrid, and Ca-astroph networks. In the Karate network, λ =(0.05-0.07) and (0.09-0.1) intervals, the Kendall ɩ value of the KLN algorithm is the highest, while λ =(0.01-0.04) The initial probability is lower than that of DC and GSI algorithms. λ = 0.08, the EC algorithm works best. In the Football network, KLN and ECRM are at a relatively high level overall, and the PL algorithm is λ =0.03 has the highest ɩ value, and the Kshell algorithm has the lowest value in this network because the network structure is relatively simple and a large number of nodes have the same Kshell value, resulting in a decrease in node differentiation. λ =(0.02-0.08), the ɩ value of KLN is at the highest level. λ The increase effect of is only lower than that of MCGM and ECRM algorithms. In the Friendships network, λ =(0.02-0.04), the performance of MCGM, GSI, and ECRM algorithms is better, and the ɩ values of KLN in other intervals are all at a high level. λ>0.8 In the intervals between the two networks, KLN achieved the highest ɩ values, and overall ranked behind ECRM and MCGM in all other intervals. In summary, the KLN algorithm achieved the best Kendall ɩ values overall in five of the ten networks, and also achieved high ɩ values in some intervals in the remaining five networks. Therefore, the KLN algorithm demonstrated high accuracy in identifying key nodes using the SIR model.
[0060] The infectious ability of all nodes in the network: Here we analyze the infectious ability of nodes in the network, map the ranking distribution of nodes in each algorithm result to the ranking result of the SIR model, and draw a curve corresponding to the number of nodes infected in the SIR model. As the importance of the nodes decreases, the ideal curve should show a smooth downward trend, proving that the ranking result of the algorithm is more accurate, and the stronger the ranking result is in terms of infectious ability in the network. In the experiment, we set λ =0.1, μ=1. The number of iterations is 1000 (except for the Ca-astroph network, which is 100). Based on the number of nodes in the network, the curve data is displayed in linear form in the first three networks and in Log10 form in the subsequent networks. Figure 7 、 Figure 8 and Figure 9 The graphs of the node infection capabilities of the algorithm of the present invention and the comparative algorithm in 10 networks are shown as follows: Figure 7 (a) Karate, (b) Dolphins, (c) Football, (d) Euroroad, and Figure 8 (e) Usair, (f) Friendships, (g) Protein, (h) Powergrid, and Figure 9 (i) HepPh and (j) Ca-astroph.
[0061] Depend on Figure 7 、 Figure 8 and Figure 9 As can be seen in the figure, the KLN algorithm exhibits a smooth downward trend in the seven networks (Karate, Dolphins, Euroroad, Friendships, Powergrid, HepPh, and Ca-astroph), achieving the best overall infection performance. In the Football network, the curves of all algorithms show little difference, due to the large variation in node degrees within the network. In the Usair network, the KLN, ECRM, and GSI algorithms achieve similar overall performance, significantly outperforming the other algorithms in infection performance. In the Protein network, the ECRM algorithm achieves the best overall infection performance, but the overall trend of the KLN algorithm's top ten nodes is similar to that of the ECRM algorithm. In summary, the KLN algorithm exhibits a smooth downward trend in infection performance across most networks, demonstrating that using this method's ranking results to infect networks is generally effective.
[0062] In order to better verify the effectiveness of the algorithm, this paper selects the first ten key nodes in the 12 algorithms to infect the surrounding nodes, sets the infection probability to 0.01 and the recovery probability to 0.1, counts the total number of infected nodes after each round, iterates 1000 times and takes the average value of the results, and counts a total of 30 rounds.
[0063] like Figure 10 、 Figure 11 As shown, Figure 10 (a) Karate, (b) Dolphins, (c) Football, (d) Euroroad, and Figure 11Comparison results for Usair (e), Friendships (f), Protein (g), Powergrid (h), HepPh (i), and Ca-astroph (j) show that the number of infected nodes for the 12 algorithms in the 10 networks increases with each round, and the rate of increase gradually slows. KLN achieves the best overall infection performance in the Euroroad, Usair, Friendships, and Protein networks. In the Karate, Dolphins, Powegrid, and CA-astroph networks, KLN approaches the optimal performance of the compared algorithms. In the Hepph network, the overall performance of the 12 compared algorithms is relatively consistent, with little differentiation. In the Football network, KLN performs in the middle, while Enrenew performs better. This is because the Enrenew algorithm suppresses neighboring nodes after selecting key nodes. Therefore, it is more suitable for selecting the top-N key nodes in clustered social networks, but is not suitable for ranking nodes across the entire network. Comprehensive analysis shows that the Top-10 nodes in the KLN algorithm can show good infection effects in most networks.
[0064] The present invention is aimed at the existing key node identification algorithms. Most of the algorithms perform multi-attribute fusion analysis of nodes under the condition of fixed network topology to achieve the distinction of node importance. They ignore the impact of node failure in real networks on changes in local network structure and do not fit the actual network environment. Therefore, the present invention proposes a novel key node identification algorithm KLN algorithm. First, by removing nodes from the network and taking advantage of the damage caused by nodes to the local network structure, from the perspective of information transmission between nodes, the probability distribution change of communication between nodes in the local network structure before and after the node removal is analyzed, and the idea of KL divergence is used to measure the information entropy lost before and after the local network change. Secondly, the KLN algorithm combines the attributes of the node itself in the global network and its own local attributes to evaluate the destructive impact of the node on the network after removal. Finally, by calculating the destructive impact of all nodes in the local network structure centered on the node, the final judgment of the node importance is achieved. Through comparative experiments with 11 comparison algorithms in 10 networks, we analyzed four aspects: the correlation of the top ten important nodes, Kendall's τ, and the infectious power of the entire network and the top ten important nodes. The results show that KLN outperforms traditional classic algorithms such as DC, EC, and Kshell. Furthermore, compared with new algorithms from recent years, such as PR, PL, Enrenew, MCDE, ECRM, GSM, GSI, and MECG, it outperforms the comparison algorithms in most cases, demonstrating better results and accuracy.
[0065] The present invention combines the characteristics of the actual transportation network environment, such as being volatile and easily destroyed, and analyzes and identifies key nodes in the network by simulating the destruction of the network structure. It can be well applied to transportation networks of most sizes and other types of complex network structures, improve the accuracy and effectiveness of important nodes, and provide effective decision support for route planning, traffic flow distribution, etc. of the transportation network.
[0066] Based on the same inventive concept, the present invention also proposes a key node identification system for a transportation network, comprising: The acquisition module is used to obtain all nodes in the transportation network structure and their Kshell values and degree values, and select one of the nodes.
[0067] The probability distribution calculation module is used to calculate the true probability distribution of information transmission between the surrounding neighboring nodes of the selected node and the selected node based on the degree values of all nodes, as well as the fitted probability distribution of information transmission between the one-hop neighboring nodes in the neighborhood after removing the edges connecting the selected node and the surrounding neighboring nodes.
[0068] The importance value calculation module is used to obtain the information entropy lost by each neighbor node of the selected node using the KL divergence model based on the true probability distribution and the fitted probability distribution; based on the information entropy lost by the neighbor node and the Kshell value, the KL divergence value - the information entropy lost caused by the removal of the selected node on all one-hop neighbor nodes in the neighborhood is obtained; based on the KL divergence value - the information loss entropy and the degree value of the selected node itself, the destructive impact of the selected node is obtained; based on the destructive impact of the selected node and the one-hop neighbor nodes in the neighborhood, the importance value of the selected node is obtained.
[0069] The key node identification module is used to identify the key nodes of transportation network structure damage by calculating the importance value of each node.
[0070] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for identifying key nodes in a transportation network, characterized in that: The following steps are involved: Obtain all nodes in the transportation network structure and their Kshell values and degree values, and select one of the nodes; Based on the degree values of all nodes, the true probability distribution of information transmission between the selected node’s neighboring nodes and the selected node is calculated, as well as the fitted probability distribution of information transmission between the selected node and the one-hop neighboring nodes in the neighborhood after removing the edges connecting the selected node and the surrounding neighboring nodes. According to the true probability distribution and the fitted probability distribution, the KL divergence model is used to obtain the information entropy lost by each neighboring node of the selected node. According to the information entropy lost by the neighboring node and the Kshell value, the KL divergence value - the information entropy lost by all one-hop neighboring nodes in the neighborhood after the selected node is removed is obtained. The destructive influence of the selected node is obtained based on the KL divergence value - information loss entropy and the degree value of the selected node itself; the importance value of the selected node is obtained based on the destructive influence of the selected node and its one-hop neighboring nodes in the neighborhood; By calculating the importance value of each node, the key nodes of transportation network structure damage are identified.
2. A method for identifying key nodes in a transportation network according to claim 1, characterized in that: The true probability distribution of neighboring nodes of the selected node transmitting information through the selected node is expressed as: ; in, n Represents the node number in the network; p i ( j ) represents the selected node i No. j The probability of information transmission between neighbor nodes; j =1,2,3... n is a node i Neighbor nodes of x is a random variable; the probability of information transmission p i ( j ) is equal to the neighbor node j The degree value and node i The proportion of the sum of the degree values of all neighbor nodes, which is expressed as: ; in, d(j) Representative neighbor node j The degree value, Γ ( i ) represents a node i The set of all one-hop neighbor nodes.
3. A method for identifying key nodes in a transportation network according to claim 2, characterized in that: After removing the edges connecting the selected node and its neighboring nodes, the fitted probability distribution of information transmission among one-hop neighboring nodes in the neighborhood is expressed as: ; in, in Representative Node i No. n neighbor nodes, q in ( x ) represents neighbor nodes n Respectively with nodes i The fitting probability distribution of information transmission among all one-hop neighbor nodes; q in ( x ) is calculated as: ; in: ; in, Representative Node i After removal, its neighbor nodes n With neighboring nodes j Fitting probability distribution for information transmission, Calculated based on the shortest distance and degree between nodes, Representative Node n and j The influence coefficient of the information transmission distance between Representative Node i After the destruction, its neighbor nodes j The current degree value.
4. A method for identifying key nodes in a transportation network according to claim 3, characterized in that: According to the true probability distribution and the fitted probability distribution, the KL divergence model is used to obtain the information entropy lost by each neighbor node of the selected node. The calculation formula is: ; in, x is a random variable, X is a collection of random variables.
5. A method for identifying key nodes in a transportation network according to claim 4, characterized in that: The KL divergence value generated by the selected node after removal for all one-hop neighbor nodes in the neighborhood - the loss information entropy D is obtained based on the information entropy and Kshell value of the neighbor node loss. i The formula for calculating (P||Q) is: 。 6. A method for identifying key nodes in a transportation network according to claim 5, characterized in that: The destructive influence of the selected node is obtained according to the KL divergence value-information loss entropy and the degree value of the selected node itself. DE ( i ), which is calculated as follows: ; in, d ( i ) represents a node i The degree value.
7. A method for identifying key nodes in a transportation network according to claim 6, characterized in that: The importance value of the selected node is obtained based on the destructive influence of the selected node and the one-hop neighbor node in the neighborhood. KLN ( i ), which is calculated as follows: 。 8. A key node identification system for a transportation network, characterized in that: include: The acquisition module is used to obtain all nodes in the transportation network structure and their Kshell values and degree values, and select one of the nodes; The probability distribution calculation module is used to calculate the true probability distribution of information transmission between the neighboring nodes of the selected node and the selected node based on the degree values of all nodes, and the fitted probability distribution of information transmission between the one-hop neighboring nodes in the neighborhood after removing the edges connecting the selected node and the surrounding neighboring nodes; The importance value calculation module is used to obtain the information entropy lost by each neighboring node of the selected node using the KL divergence model based on the true probability distribution and the fitted probability distribution; based on the information entropy lost by the neighboring node and the Kshell value, the KL divergence value - the information entropy lost by all one-hop neighboring nodes in the neighborhood after the selected node is removed is obtained; The destructive influence of the selected node is obtained based on the KL divergence value - information loss entropy and the degree value of the selected node itself; the importance value of the selected node is obtained based on the destructive influence of the selected node and its one-hop neighboring nodes in the neighborhood; The key node identification module is used to identify the key nodes of transportation network structure damage by calculating the importance value of each node.