Bentweenness centrality estimation method and device based on sampling strategy

By adopting the internumeric centrality estimation method based on sampling strategy in large-scale networks, the problem that traditional computing methods are difficult to meet the computing needs of large-scale networks is solved, and the calculation speed is improved and error control is achieved.

CN120223573AActive Publication Date: 2025-06-27NAT UNIV OF DEFENSE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510441455.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-27
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Traditional interdigital-centric computing methods are difficult to meet the computing needs in large-scale networks, and researchers hope to speed up the computing speed by reducing the calculation accuracy.

Method used

The median centrality estimation method based on the sampling strategy is used to process the network to obtain the centrality information of the network nodes, and then process this information, select a selected node set, and finally calculate the median centrality approximation.

Benefits of technology

This method can significantly speed up the calculation speed while reducing calculation errors, and is suitable for internumerical central calculations of large-scale networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223573A_ABST
    Figure CN120223573A_ABST
Patent Text Reader

Abstract

The invention discloses a betweenness centrality estimation method and device based on a sampling strategy, and the method comprises the steps: carrying out the processing of a network G, and obtaining the centrality information of a network node; processing the centrality information of the network nodes to obtain a selected node set S; and processing the selected node set S to obtain a betweenness centrality approximate value. The method is low in calculation error, more obvious in error rule, more stable in change and better in effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network technologies, and in particular, to a method and device for estimating betweenness centrality based on a sampling strategy. Background Art

[0002] Since the 21st century, the informatization of human society has developed rapidly. The popularization and application of mobile Internet and the rapid implementation of Internet of Things deployment have significantly increased the number of devices on the network, and the complexity of the number of physical network devices and connection relationships has been increasing day by day. Real networks in different fields such as power grids, telecommunications networks, transportation networks, social networks, communication networks, and Internet of Things have formed an all-inclusive network world.

[0003] By abstracting different real networks, the corresponding abstract networks often exhibit the same network properties. The systematic study of the properties and laws of abstract networks has formed a complex network science with rich connotations. The main research topics of complex network science cover network centrality measurement and global network characteristics, network model structure and function analysis, network link prediction and recommendation algorithms, network dynamics, network control and optimization, etc. The research on network centrality measurement plays a fundamental role in complex network science. In order to characterize the importance of nodes and edges in a network, researchers have proposed a variety of network centrality measures to evaluate the centrality of nodes and edges in a network from multiple perspectives. Network centrality measurement mainly characterizes the importance of nodes and edges in a network, providing theoretical support for the identification of key nodes and key edges in real networks.

[0004] Betweenness centrality is an index that characterizes the importance of a certain node or an edge in a network from the perspective of network connectivity, and it is also an important research point in complex networks. The research on its fast calculation method has far-reaching practical significance. For example, according to the core nodes in a social network, an influence ranking list is given to attract users and conduct targeted network marketing; by controlling the leader of a terrorist organization, the entire terrorist organization can be controlled, thereby avoiding terrorist attack cases; by protecting the key servers in a network, the key servers can be prevented from being attacked by viruses or hackers, so that the entire network can operate normally; by isolating the source of infection, the spread and diffusion of infectious disease viruses can be effectively prevented, etc. In these practical applications, it is necessary to know the importance degree of each node in the network in order to find out the key nodes in the network.

[0005] However, with the growth of the number of nodes in the network and the denseness of the connection relationships between nodes, the topological structure of the network becomes more and more complex. The traditional calculation method of betweenness centrality is difficult to fully meet the requirements of calculating betweenness centrality in large-scale networks. Researchers hope to speed up the calculation speed by reducing part of the calculation accuracy to calculate the betweenness centrality in large-scale networks. Therefore, the approximate calculation of betweenness centrality has become one of the hot issues in the field of complex network analysis. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a betweenness centrality estimation method and device based on a sampling strategy, which can accelerate the calculation speed by reducing the calculation precision of some parts, so as to calculate the betweenness centrality in a large-scale network.

[0007] To solve the above technical problem, a first aspect of an embodiment of the present invention discloses a betweenness centrality estimation method based on a sampling strategy, and the method includes:

[0008] S1, process the network G to obtain the centrality information of network nodes;

[0009] S2, process the centrality information of the network nodes to obtain a selected node set S;

[0010] S3, process the selected node set S to obtain an approximate value of betweenness centrality.

[0011] As an optional implementation manner, in the first aspect of an embodiment of the present invention, the process of processing the network G to obtain the centrality information of network nodes includes:

[0012] S11, process the network G to obtain network system parameter information; the network system parameter information includes the extraction ratio p of network nodes and the termination condition;

[0013] S12, use the centrality calculation model of network nodes, and process the network G according to the network system parameter information to obtain the centrality information of network nodes;

[0014] The expression of the centrality calculation model of the network nodes is:

[0015]

[0016] where PR(i) is the centrality information of node i, PR(j) is the centrality information of node j, c is the random jump probability, n is the number of vertices in the network, is the out-degree of node j, a ij is the weight between node i and node j.

[0017] As an optional implementation manner, in the first aspect of an embodiment of the present invention, the process of processing the centrality information of the network nodes to obtain the selected node set S includes:

[0018] S21, sort the nodes of the network G according to the centrality information of the network nodes to obtain an ordered sequence of network nodes;

[0019] S22. Select the first p·n nodes from the ordered sequence of the network nodes to obtain a selected node set S.

[0020] As an alternative implementation, in the first aspect of the embodiments of the present invention, the processing of the selected node set S to obtain an approximation of betweenness centrality includes:

[0021] S31. Traverse the selected node set S to obtain the shortest paths between nodes;

[0022] S32. Process the shortest paths between nodes and the selected node set S to obtain the number of times the shortest paths between nodes pass through each vertex and edge in the network;

[0023] S33. Process the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality.

[0024] As an alternative implementation, in the first aspect of the embodiments of the present invention, the processing of the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality includes:

[0025] Use an approximation calculation model of betweenness centrality to process the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality;

[0026] The expression of the approximation calculation model of betweenness centrality is:

[0027]

[0028] where σ(v i ,v j ) is the number of shortest paths from node v i to node v j , σ(v i ,v j |v) is the number of shortest paths from node v i to node v j passing through node v, V is the vertex set of the network G, and Bet(v) is the approximation of betweenness centrality of node v.

[0029] As an alternative implementation, in the first aspect of the embodiments of the present invention, the expression of the network G is:

[0030] G=(V,E)

[0031] Among them, V represents the vertex set of the network, E represents the edge set of the network, n = |V| represents the number of vertices in the network, and m = |E| represents the number of edges in the network.

[0032] In the second aspect of the embodiments of the present invention, an apparatus for estimating betweenness centrality based on a sampling strategy is disclosed. The apparatus includes:

[0033] A centrality information calculation module, configured to process the network G to obtain the centrality information of network nodes;

[0034] A selected node set calculation module, configured to process the centrality information of the network nodes to obtain a selected node set S;

[0035] A betweenness centrality approximation calculation module, configured to process the selected node set S to obtain an approximation of the betweenness centrality.

[0036] As an optional implementation manner, in the second aspect of the embodiments of the present invention, the processing of the network G to obtain the centrality information of network nodes includes:

[0037] S11, processing the network G to obtain network system parameter information; the network system parameter information includes the extraction ratio p of network nodes and the termination condition;

[0038] S12, using the centrality calculation model of network nodes, and processing the network G according to the network system parameter information to obtain the centrality information of network nodes;

[0039] The expression of the centrality calculation model of the network nodes is:

[0040]

[0041] Among them, PR(i) is the centrality information of node i, PR(j) is the centrality information of node j, c is the random jump probability, n is the number of vertices in the network, is the out-degree of node j, a ij is the weight between node i and node j.

[0042] As an optional implementation manner, in the second aspect of the embodiments of the present invention, the processing of the centrality information of the network nodes to obtain the selected node set S includes:

[0043] S21, sorting the nodes of the network G according to the centrality information of the network nodes to obtain an ordered sequence of network nodes;

[0044] S22, selecting the first p·n nodes from the ordered sequence of network nodes to obtain the selected node set S.

[0045] As an alternative implementation, in the second aspect of the embodiments of the present invention, the processing of the selected node set S to obtain an approximation of betweenness centrality includes:

[0046] S31, traversing the selected node set S to obtain the shortest paths between nodes;

[0047] S32, processing the shortest paths between nodes and the selected node set S to obtain the number of times the shortest paths between nodes pass through each vertex and edge in the network;

[0048] S33, processing the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality.

[0049] As an alternative implementation, in the second aspect of the embodiments of the present invention, the processing of the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality includes:

[0050] Using an approximation calculation model of betweenness centrality to process the shortest paths between nodes and the number of times the shortest paths between nodes pass through each vertex and edge in the network to obtain an approximation of betweenness centrality;

[0051] The expression of the approximation calculation model of betweenness centrality is:

[0052]

[0053] where σ(v i ,v j ) is the number of shortest paths from node v i to node v j , σ(v i ,v j |v) is the number of shortest paths from node v i to node v j passing through node v, V is the vertex set of network G, and Bet(v) is the approximation of betweenness centrality of node v.

[0054] As an alternative implementation, in the second aspect of the embodiments of the present invention, the expression of network G is:

[0055] G=(V,E)

[0056] where V represents the vertex set of the network, E represents the edge set of the network, n = |V| represents the number of vertices in the network, and m = |E| represents the number of edges in the network.

[0057] The third aspect of the present invention discloses another betweenness centrality estimation device based on a sampling strategy, and the device includes:

[0058] A memory storing executable program code;

[0059] A processor coupled to the memory;

[0060] The processor calls the executable program code stored in the memory and executes some or all of the steps in the betweenness centrality estimation method based on a sampling strategy disclosed in the first aspect of the embodiments of the present invention.

[0061] The fourth aspect of the present invention discloses a computer-readable storage medium storing computer instructions, which are used to execute some or all of the steps in the betweenness centrality estimation method based on a sampling strategy disclosed in the first aspect of the embodiments of the present invention when the computer instructions are called.

[0062] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0063] The present invention discloses a betweenness centrality estimation method and device based on a sampling strategy. By processing the network G, the centrality information of network nodes is obtained; by processing the centrality information of network nodes, a selected node set S is obtained; by processing the selected node set S, an approximate value of betweenness centrality is obtained. The method of the present invention has a lower calculation error, a more obvious error pattern, more stable changes, and better effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 is a flowchart of a betweenness centrality estimation method based on a sampling strategy disclosed in the embodiments of the present invention;

[0066] Figure 2 is a flowchart of another betweenness centrality estimation method based on a sampling strategy disclosed in the embodiments of the present invention;

[0067] Figure 3 is the sampling error of random network based on PageRank centrality sampling in different cases disclosed in the embodiments of the present invention;

[0068] Figure 4Sampling errors of scale-free networks based on PageRank centrality sampling in different cases disclosed in the embodiments of the present invention;

[0069] Figure 5 Sampling errors of small-world networks based on PageRank centrality sampling in different cases disclosed in the embodiments of the present invention;

[0070] Figure 6 Schematic structural diagram of a betweenness centrality estimation device based on a sampling strategy disclosed in the embodiments of the present invention;

[0071] Figure 7 Schematic structural diagram of another betweenness centrality estimation device based on a sampling strategy disclosed in the embodiments of the present invention. Specific implementation manners

[0072] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0073] The terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or equipment.

[0074] Referring to "embodiments" herein means that a specific feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0075] The present invention discloses a method and apparatus for estimating betweenness centrality based on a sampling strategy. The method includes: processing network G to obtain centrality information of network nodes; processing the centrality information of the network nodes to obtain a selected node set S; and processing the selected node set S to obtain an approximate value of betweenness centrality. The method of the present invention has a lower calculation error, a more obvious error pattern, more stable changes, and better effects. The following will be described in detail respectively.

[0076] Embodiment 1

[0077] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for estimating betweenness centrality based on a sampling strategy disclosed in an embodiment of the present invention. Among them, Figure 1 the described method for estimating betweenness centrality based on a sampling strategy is applied to the field of network technology, and the embodiment of the present invention does not make any limitations. As Figure 1 shown, the method for estimating betweenness centrality based on a sampling strategy may include the following operations:

[0078] S1. Process network G to obtain centrality information of network nodes;

[0079] S2. Process the centrality information of the network nodes to obtain a selected node set S;

[0080] S3. Process the selected node set S to obtain an approximate value of betweenness centrality.

[0081] Optionally, the processing network G to obtain centrality information of network nodes includes:

[0082] S11. Process network G to obtain network system parameter information; the network system parameter information includes the extraction ratio p of network nodes and the termination condition;

[0083] S12. Use the centrality calculation model of network nodes, and process network G according to the network system parameter information to obtain centrality information of network nodes;

[0084] The expression of the centrality calculation model of network nodes is:

[0085]

[0086] where PR(i) is the centrality information of node i, PR(j) is the centrality information of node j, c is the random jump probability, n is the number of vertices in the network, is the out-degree of node j, and a ij is the weight between node i and node j.

[0087] Optionally, another implementation method of step S12 is as follows:

[0088] 1) Read the network G, G = (V, E), V = {v1, v2, …, v n}, the elements in V are the nodes (also called vertices) of the network G, n is the number of nodes, and E = {e1, e2, …, e m} is the edge set, m is the number of edges, and the element e i in E = {v i1 , v i2 , …, v ij} is a hyperedge, where i = 1, 2, …, m and j = 1, 2, …, n;

[0089] 2) Process the network G to obtain the incidence matrix C. C is an n×m matrix, and the elements c ij in the matrix take 0 or 1. When the node v i is located on the hyperedge e j , c ij = 1;

[0090] 3) Transform H to obtain the adjacency matrix A. The element a ij in A is the number of nodes v i and node v j that are located on the same hyperedge;

[0091] The transformation method is as follows:

[0092] Input the matrix H

[0093] Input the transformation option (0: adjacency → incidence; 1: incidence → adjacency)

[0094] If opt = 1

[0095] n = the number of rows of H (number of nodes)

[0096] m = the number of columns of H (number of hyperedges)

[0097] Initialize the adjacency matrix, M1 = Zeros(n, n)

[0098] for i = 1 to m:

[0099] a = the row information where each column in the H matrix has non - zero terms

[0100] If a is not empty

[0101] all connected = the direct product (Cartesian product) of the row information of a

[0102] for traversing all connected

[0103] If the coordinate is not on the main diagonal of M

[0104] Write: increment the pre-written coordinate value by 1 (auto-increment by 1)

[0105] 4) Calculate the centrality information of network nodes:

[0106]

[0107] where PR is the centrality information of network nodes and M is the random transition matrix:

[0108]

[0109] D v is the out-degree of the node, d H (j) is the out-degree of node v i with respect to the hyper-edge e j and H(i) is the set of out-degrees of node v i in the matrix H, D e is the adjacency degree of the node, d A (j) is the adjacency degree of node v i with respect to node v j and A(i) is the set of adjacency degrees of node v i in the matrix A; the matrix a = [a0, a1,..., a n is the correction matrix. When a column is all 0 (the node is an isolated node), a i = 1, i = 0, 1,..., n, and for the others a i = 0, e is the identity matrix, β represents the random walk probability of the node, β takes values from 0.8 to 0.9, and the specific value is set by experiments, W is the node similarity matrix, and the element a(i) is the set of adjacent nodes of node v i and a(j) is the set of adjacent nodes of node v j .

[0110] Optionally, the processing of the centrality information of the network nodes to obtain the selected node set S includes:

[0111] S21. Sort the nodes of network G according to the centrality information of the network nodes to obtain an ordered sequence of network nodes;

[0112] S22. Select the first p·n nodes from the ordered sequence of network nodes to obtain the selected node set S.

[0113] Optionally, the processing of the selected node set S to obtain an approximate value of betweenness centrality includes:

[0114] S31. Traverse the selected node set S to obtain the shortest paths between nodes;

[0115] S32. Process the shortest paths between nodes and the selected node set S to obtain the number of times each vertex and edge in the network is traversed by the shortest paths between nodes;

[0116] S33. Process the shortest paths between nodes and the number of times each vertex and edge in the network is traversed by the shortest paths between nodes to obtain an approximation of betweenness centrality.

[0117] Optionally, the process of processing the shortest paths between nodes and the number of times each vertex and edge in the network is traversed by the shortest paths between nodes to obtain an approximation of betweenness centrality includes:

[0118] Use an approximation calculation model of betweenness centrality to process the shortest paths between nodes and the number of times each vertex and edge in the network is traversed by the shortest paths between nodes to obtain an approximation of betweenness centrality;

[0119] The expression of the approximation calculation model of betweenness centrality is:

[0120]

[0121] where σ(v i ,v j ) is the number of shortest paths from node v i to node v j , σ(v i ,v j |v) is the number of shortest paths from node v i to node v j that pass through node v, V is the vertex set of network G, and Bet(v) is the approximation of betweenness centrality of node v.

[0122] Optionally, the expression of network G is:

[0123] G=(V,E)

[0124] where V represents the vertex set of the network, E represents the edge set of the network, n = |V| represents the number of vertices in the network, and m = |E| represents the number of edges in the network.

[0125] It can be seen that the present invention discloses a method and device for estimating betweenness centrality based on a sampling strategy. By processing network G, the centrality information of network nodes is obtained; by processing the centrality information of network nodes, a selected node set S is obtained; by processing the selected node set S, an approximation of betweenness centrality is obtained. The method of the present invention has a lower calculation error, a more obvious error pattern, a more stable change, and a better effect.

[0126] Example 2

[0127] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of another betweenness centrality estimation method based on a sampling strategy disclosed in an embodiment of the present invention. Among them, Figure 2 The described betweenness centrality estimation method based on a sampling strategy is applied to the field of network technology, and the embodiments of the present invention do not make any limitations. As Figure 2 shown, the betweenness centrality estimation method based on a sampling strategy may include the following operations:

[0128] 1. System parameter design

[0129] For a given network G, the user sets system parameters such as the extraction ratio p (0 ≤ p ≤ 1) of network nodes, termination conditions, etc., according to the network scale, connection relationship, and degree distribution characteristics.

[0130] 2. PageRank centrality calculation

[0131] For a given network G, the user runs an efficient PageRank calculation method to obtain the PageRank centrality values of each node in the network, and obtains an ordered set of network nodes according to the PageRank centrality values.

[0132] 3. Selection of a set of selected nodes

[0133] According to the order of the PageRank centrality values of network nodes, the first p·n nodes in the node set are selected as the set of selected nodes S, and the nodes in the set S are all important representative nodes in the network.

[0134] 4. Betweenness centrality calculation

[0135] Given the selected set S, taking the set S as an approximation of the entire vertex set V, traverse the shortest paths between the nodes in the selected node set and calculate the number of times the shortest paths pass through each vertex and edge in the network, and obtain an approximate value of the betweenness centrality using the shortest paths between the nodes in the selected node set. Its exact mathematical meaning can be expressed by the formula:

[0136]

[0137] where σ(v i ,v j ) is the number of shortest paths from node v i to node v j , and σ(v i ,v j |v) is the number of shortest paths from node v i to node v jThe number of the shortest paths passing through node v, where V is the set of vertices of network G and Bet(v) is the approximate value of the betweenness centrality of node v.

[0138] It can be seen that the present invention discloses a method and device for estimating betweenness centrality based on a sampling strategy. By processing network G, the centrality information of network nodes is obtained; by processing the centrality information of network nodes, a selected node set S is obtained; by processing the selected node set S, an approximate value of betweenness centrality is obtained. The method of the present invention has a lower calculation error, a more obvious error pattern, more stable changes, and better effects.

[0139] Embodiment III

[0140] This embodiment is carried out under three different types of networks, namely a random network, a scale-free network, and a small-world network. The generation of these three networks can all be implemented by programming:

[0141] ① Random network

[0142] A random network refers to a network in which the connection probability between all nodes is formed according to a preset value. The Erdos-Renyi network is a typical random network. In the Erdos-Renyi random network, the number of nodes n in the network usually remains unchanged. Two nodes are randomly selected to form an edge according to a fixed probability value p, then the number of edges in the network belongs to a random variable, and its expected value is:

[0143]

[0144] Later, researchers found that many important characteristics of the Erdos-Renyi random network appear instantaneously. However, with the continuous in-depth study of the Erdos-Renyi network, researchers found that the Erdos-Renyi random network is very different from other model networks. First, it does not have a high clustering coefficient; second, its degree distribution follows a Poisson distribution, rather than the scale-free characteristics that usually exist in the real world.

[0145] G = nx.fast_gnp_random_graph(num_nodes, 0.02) # Generate a random network containing num_nodes nodes with an edge probability of 0.02

[0146] ② Scale-free network

[0147] A class of non-uniform networks that are relatively common in the real world, such as the World Wide Web, the human metabolic network, etc. Their degree distribution is different from the uniform networks studied in the past. Their degree distribution follows a power-law characteristic, that is, a small number of nodes in the network hold the vast majority of the connected edges in the network, while a large number of nodes in the network only contain a small number of connected edges. The degree distribution corresponding to these non-uniform networks, that is, the power-law distribution (p(k) ∝ k-γ), is called a scale-free network for non-uniform networks that conform to the power-law distribution.

[0148] G = nx.barabasi_albert_graph(num_nodes, int(num_nodes * 0.001), seed = 42) # Generate a scale-free network with num_nodes nodes, each new node connecting to num_nodes * 0.001 edges, and 5 initial nodes

[0149] ③ Small-world network

[0150] Small-world networks have good clustering characteristics, similar to regular networks. However, the average path length of small-world networks is similar to that of random networks. Below, the Newman-Watts small-world model is described in detail:

[0151] Starting from a regular network: For a nearest-neighbor coupled network with n nodes, all nodes in the network are connected to s ∈ N neighbor nodes on both the left and right sides, where s is an integer;

[0152] Application of the random rewiring mechanism: Randomly select a pair of unconnected nodes for edge connection. There is only one edge between any two different nodes and it is not connected to itself.

[0153] G = nx.watts_strogatz_graph(num_nodes, int(num_nodes * 0.01), 0.2, seed = 42) # Generate a small-world network with num_nodes nodes, each node connecting to num_nodes * 0.01 neighbors, and a rewiring probability of 0.2

[0154] 1. Sampling based on PageRank centrality

[0155] First, select a network model to generate and calculate the PageRank centrality of all nodes in the network. Then assign the value of PageRank centrality to the nodes as the probability of being selected in the sampling. The larger the degree, the greater the probability of being selected. Finally, according to the input sampling rate, extract the corresponding number of nodes and calculate the betweenness centrality of the selected nodes. The core code of this algorithm is:

[0156]

[0157]

[0158] 2. Experimental Efficiency Analysis

[0159] In this part, we conducted sampling experiments based on random networks, scale-free networks, and small-world networks respectively, and analyzed the experimental data from two aspects: computational accuracy and time complexity.

[0160] (1) Random Network

[0161] In the random network, first, we generated random networks with 1000 and 2000 nodes respectively, and edge probabilities of 0.1, 0.06, and 0.02. Then, in the networks with 1000 and 2000 nodes, we sampled the corresponding number of nodes at sampling rates of 0.1, 0.2, 0.3, and 0.5 respectively for calculation. The errors between the betweenness centrality of the nodes calculated based on various sampling methods and the exactly calculated betweenness centrality of the nodes are shown in the following table:

[0162] We used the NetworkX database to generate random networks with 1000 and 2000 nodes respectively. Each random network has three edge probabilities: 0.1, 0.06, and 0.02. Then, we sampled the random network based on the PageRank centrality, and calculated the betweenness centrality using the sampled nodes. The errors compared with the exactly calculated betweenness centrality of the nodes are shown in the following table:

[0163] Table 1 Sampling Error Based on PageRank Centrality

[0164]

[0165]

[0166] Table 2 Comparison of Sampling Time Based on PageRank Centrality

[0167]

[0168] It can be seen from the data that in the same network, as the sampling ratio increases, the error gradually decreases; when the sampling ratio and the edge connection probability remain unchanged, as the network scale increases, the error generally decreases; when the sampling ratio and the network scale remain unchanged, as the edge connection probability (i.e., the density) increases, the error generally decreases, but the error increases when the sampling ratio is 0.1. Generally speaking, in the 2000-node network, when the edge connection probability is 0.1 and the sampling ratio is 0.5, the error is the smallest, indicating that based on the sampling strategy of PageRank centrality in a random network, the larger the network scale, the denser the network, and the higher the sampling rate, the better the effect. Figure 3It is the sampling error of PageRank centrality sampling in different situations disclosed in the embodiments of the present invention.

[0169] (2) Scale-free network

[0170] In the scale-free network, first, we generated scale-free networks with 1000 and 2000 nodes respectively, and the edge connection quantity ratios were 0.01, 0.03, and 0.05 respectively. Then, in the networks with 1000 and 2000 nodes, corresponding numbers of nodes were sampled at sampling rates of 0.1, 0.2, 0.3, and 0.5 respectively for calculation. The errors between the node betweenness centrality calculated based on various sampling methods and the accurately calculated node betweenness centrality are shown in the following table:

[0171] First, scale-free networks with 1000 and 2000 nodes were generated using the NetworkX database, and each scale-free network had three edge connection quantity ratios of 0.01, 0.03, and 0.05. Then, the scale-free network was sampled based on PageRank centrality, and the betweenness centrality of the sampled nodes was calculated. The errors compared with the accurately calculated node betweenness centrality are shown in the following table:

[0172] Table 3 Sampling error based on PageRank centrality

[0173]

[0174] Table 4 Comparison of sampling time based on PageRank centrality

[0175]

[0176] It can be seen from the data that in the same network, as the sampling ratio increases, the error gradually decreases; when the sampling ratio and the edge connection quantity remain unchanged, as the network scale increases, the error generally decreases; when the sampling ratio and the network scale remain unchanged, as the edge connection quantity, that is, the density, increases, the error generally decreases, but the error is unstable when the sampling ratio is 0.1. Generally speaking, when the network has 2000 nodes, the edge connection probability is 0.1, and the sampling ratio is 0.5, the error is the smallest, indicating that based on the PageRank centrality sampling strategy in the scale-free network, the larger the network scale, the denser the network, and the higher the sampling rate, the better the effect. Figure 4 It is the sampling error of PageRank centrality sampling in different situations disclosed in the embodiments of the present invention;

[0177] (3) Small-world network

[0178] In the small-world network, first, we generated small-world networks with 1000 and 2000 nodes respectively, and the node neighbor ratios were 0.03, 0.06, and 0.09 respectively. Then, in the networks with 1000 and 2000 nodes, the corresponding number of nodes were sampled at sampling rates of 0.1, 0.2, 0.3, and 0.5 respectively for calculation. The errors between the betweenness centrality of nodes calculated based on various sampling methods and the exactly calculated betweenness centrality of nodes are shown in the following table:

[0179] We used the NetworkX database to generate small-world networks with 1000 and 2000 nodes respectively. Each small-world network had three node neighbor ratios: 0.03, 0.06, and 0.09. Then, based on the PageRank centrality, the small-world networks were sampled, and the betweenness centrality was calculated using the sampled nodes. The errors from the exactly calculated betweenness centrality of nodes are shown in the following table:

[0180] Table 5 Sampling errors based on PageRank centrality

[0181]

[0182] Table 6 Comparison of sampling times based on PageRank centrality

[0183]

[0184] It can be seen from the data that in the same network, as the sampling ratio increases, the errors generally decrease except for the sampling probability of 0.1; when the sampling ratio and the neighbor quantity ratio remain unchanged, as the network scale increases, the errors generally decrease; when the sampling ratio and the network scale remain unchanged, as the neighbor quantity ratio (i.e., the density) increases, the errors are generally unstable. Generally speaking, for the 1000-node network, when the neighbor quantity ratio is 0.03 and the sampling ratio is 0.2, the error is the smallest, indicating that based on the sampling strategy of PageRank centrality in the small-world network, the larger the network scale and the higher the sampling rate, the better the effect. Figure 5 It is the sampling error of sampling based on PageRank centrality disclosed in the embodiments of the present invention under different circumstances;

[0185] 3. Comparative analysis of experimental results

[0186] When the number of generated network nodes is larger, the calculation error rates are generally relatively high when the sampling ratios are 10% or 20% because the larger the network scale, the more complex the structure. The sampling ratios of 10% and 20% cannot meet the sampling requirements, and the representativeness will decrease. However, when the sampling probability reaches 50%, the sampling ratio can meet the sampling requirements, and the error will not increase significantly.

[0187] In terms of running time, as the network scale increases, the network density increases, and the sampling ratio increases, the running time will increase to a certain extent.

[0188] At the same time, it can be seen that the calculation error of the sampling strategy based on PageRank centrality is slightly lower than that of the sampling strategy based on eigenvector centrality, and the law of the calculation error of the sampling strategy based on PageRank centrality is more obvious, the change is more stable, and the effect is better.

[0189] Example Four

[0190] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a betweenness centrality estimation device based on a sampling strategy disclosed in an embodiment of the present invention. Among them, Figure 6 The described betweenness centrality estimation device based on a sampling strategy is applied to the field of network technology, and the embodiments of the present invention do not make limitations. As Figure 6 shown, the betweenness centrality estimation device based on a sampling strategy may include the following operations:

[0191] S301, a centrality information calculation module, configured to process the network G to obtain the centrality information of network nodes;

[0192] S302, a selected node set calculation module, configured to process the centrality information of the network nodes to obtain a selected node set S;

[0193] S303, a betweenness centrality approximation calculation module, configured to process the selected node set S to obtain an approximation of the betweenness centrality.

[0194] Example Five

[0195] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of another betweenness centrality estimation device based on a sampling strategy disclosed in an embodiment of the present invention. Among them, Figure 7 The described betweenness centrality estimation device based on a sampling strategy is applied to the field of network technology, and the embodiments of the present invention do not make limitations. As Figure 7 shown, the betweenness centrality estimation device based on a sampling strategy may include the following:

[0196] A memory 401 storing executable program code;

[0197] A processor 402 coupled to the memory 401;

[0198] The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the betweenness centrality estimation method based on a sampling strategy described in Example One, Example Two, and Example Three.

[0199] Example VI

[0200] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program causes a computer to execute the steps in the betweenness centrality estimation method based on a sampling strategy described in Embodiment I, Embodiment II, and Embodiment III.

[0201] The device embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separated, and components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0202] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each implementation can be realized by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0203] Finally, it should be noted that: What is disclosed by an eigenvector centrality estimation method and device based on a sampling strategy disclosed in the embodiments of the present invention is only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than limiting it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for estimating betweenness centrality based on sampling strategy, characterized in that: The method comprises: S1, process the network G to obtain the centrality information of the network nodes, including: S11, processing the network G to obtain network system parameter information; the network system parameter information includes the extraction ratio p of the network nodes and the termination condition; S12, using a centrality calculation model of network nodes, processing the network G according to the network system parameter information to obtain centrality information of network nodes; The centrality calculation model expression of the network node is: Among them, PR(i) is the centrality information of node i, PR(j) is the centrality information of node j, c is the random jump probability, n is the number of vertices in the network, is the out-degree of node j, a ij is the weight between node i and node j; S2, processing the centrality information of the network nodes to obtain a selected node set S, including: S21, sorting the nodes of the network G according to the centrality information of the network nodes to obtain an ordered sequence of the network nodes; S22, selecting the first p·n nodes from the ordered sequence of network nodes to obtain a selected node set S; S3, processing the selected node set S to obtain an approximate value of betweenness centrality, including: S31, traversing the selected node set S to obtain the shortest path between the nodes; S32, processing the shortest path between the nodes and the selected node set S to obtain the number of times the shortest path between the nodes passes through each vertex and edge in the network; S33, processing the shortest path between the nodes and the number of times the shortest path between the nodes passes through each vertex and edge in the network to obtain an approximate value of betweenness centrality.

2. The betweenness centrality estimation method based on sampling strategy according to claim 1 is characterized in that: The shortest path between the nodes and the number of times the shortest path between the nodes passes through each vertex and edge in the network are processed to obtain an approximate value of betweenness centrality, including: Using the betweenness centrality approximate value calculation model, the shortest path between the nodes and the number of times the shortest path between the nodes passes through each vertex and edge in the network are processed to obtain the betweenness centrality approximate value; The betweenness centrality approximate value calculation model expression is: Among them, σ(v i ,v j ) is the node v i To node v j The number of shortest paths, σ(v i ,v j |v) is the node v i To node v j The number of shortest paths passing through node v, V is the vertex set of network G, and Bet(v) is the approximate betweenness centrality of node v.

3. The betweenness centrality estimation method based on sampling strategy according to claim 1 is characterized in that: The expression of the network G is: G=(V,E) Among them, V represents the vertex set of the network, E represents the edge set of the network, n = |V| represents the number of vertices in the network, and m = |E| represents the number of edges in the network.

4. A device for estimating betweenness centrality based on sampling strategy, characterized in that: The device comprises: The centrality information calculation module is used to process the network G to obtain the centrality information of the network nodes, including: S11, processing the network G to obtain network system parameter information; the network system parameter information includes the extraction ratio p of the network nodes and the termination condition; S12, using a centrality calculation model of network nodes, processing the network G according to the network system parameter information to obtain centrality information of network nodes; The centrality calculation model expression of the network node is: Among them, PR(i) is the centrality information of node i, PR(j) is the centrality information of node j, c is the random jump probability, n is the number of vertices in the network, is the out-degree of node j, a ij is the weight between node i and node j; The selected node set calculation module is used to process the centrality information of the network nodes to obtain the selected node set S, including: S21, sorting the nodes of the network G according to the centrality information of the network nodes to obtain an ordered sequence of the network nodes; S22, selecting the first p·n nodes from the ordered sequence of network nodes to obtain a selected node set S; The betweenness centrality approximate value calculation module is used to process the selected node set S to obtain the betweenness centrality approximate value, including: S31, traversing the selected node set S to obtain the shortest path between the nodes; S32, processing the shortest path between the nodes and the selected node set S to obtain the number of times the shortest path between the nodes passes through each vertex and edge in the network; S33, processing the shortest path between the nodes and the number of times the shortest path between the nodes passes through each vertex and edge in the network to obtain an approximate value of betweenness centrality.

5. A device for estimating betweenness centrality based on sampling strategy, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the betweenness centrality estimation method based on sampling strategy as described in any one of claims 1-3.

6. A computer storable medium, characterized in that: The computer storable medium stores computer instructions, and when the computer instructions are called, they are used to execute the betweenness centrality estimation method based on sampling strategy as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • System for stratified sampling of representative sub-networks from complex network

    CN113901982A

  • Betweenness centrality approximate calculation method and device

    CN116032828A

  • Complex network vulnerability assessment method, device and system, and storage medium

    CN118839445A

  • Network node importance quantitative evaluation method based on deep learning

    CN119052108A