Social network key node mining method and mining system thereof

By combining multiple indicators for pre-screening and anti-power method to calculate the minimum eigenvalue of the deleted Laplacian matrix, the problem of rapid accuracy of key node mining in social networks is solved, and efficient information dissemination management and public opinion guidance are achieved.

CN120179920APending Publication Date: 2025-06-20HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234391.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

It is difficult for existing technology to quickly and accurately unearth key nodes in social networks, resulting in efficiency and accuracy problems in information dissemination management, public opinion guidance, marketing promotion and crisis management.

Method used

By comprehensively considering multiple indicators (degree, median, K_Shell value, resistance distance), pre-selected node groups are constructed, and then the minimum eigenvalue of the deleted Laplacian matrix is ​​calculated by using the anti-power method to further screen, and the key nodes are gradually determined.

Benefits of technology

It has achieved rapid and accurate mining of key nodes in social networks, improved the efficiency and effectiveness of tasks such as information dissemination management and public opinion guidance, and at the same time reduced the complexity of the algorithm, and is suitable for large-scale complex networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179920A_ABST
    Figure CN120179920A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field related to information retrieval, and discloses a social network key node mining method and system, and the mining method comprises the steps: S1, comprehensively considering a plurality of indexes, screening p pre-selected nodes from all nodes of a social network, and forming a pre-selected node group S, the plurality of indexes comprise at least three of degree, betweenness, KShell value and resistance distance; s2, selecting a pre-selected node which does not belong to the node group T from the S, calculating a minimum characteristic value of a Laplacian matrix after T'deletion of the social network by using an inverse power method, taking the minimum characteristic value as a score of the corresponding pre-selected node, adding the pre-selected node v with the highest score into the T, and enabling the T to be a null set at an initial moment; and S3, repeatedly executing the step S2 until the number of the nodes in the T reaches a preset scale, and taking the nodes in the T as key nodes. The method has the advantages that nodes with high importance can be accurately found out, and meanwhile it can be guaranteed that the algorithm complexity is low to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to information retrieval, and more specifically, relates to a method and a system for mining key nodes in a social network. Background Art

[0002] In recent years, social networks have become an important platform for information dissemination, and their influence has become more and more significant with the continuous expansion of the network scale. In this era of highly interconnected information, how to effectively manage and control information dissemination, especially mining and utilizing key nodes in the network, has become an urgent problem to be solved in the academic and practical fields.

[0003] Complex network theory provides us with a powerful tool to understand the mechanism of information flow. By mining key nodes, we can identify nodes that play important roles in information dissemination, and these nodes are often the "bridges" or "hubs" of information dissemination. In-depth analysis of these key nodes not only helps to reveal the internal laws of information flow, but also provides important theoretical basis and practical guidance for public opinion guidance, market promotion and crisis management in social networks. In terms of public opinion guidance, understanding the characteristics and behavior patterns of key nodes can help policymakers effectively design information dissemination strategies to actively influence public opinion and enhance social consensus. In market promotion, enterprises can formulate more targeted marketing strategies by mining and utilizing these key nodes to improve the efficiency and effectiveness of brand dissemination. At the same time, in crisis management, timely mining of key nodes in information dissemination helps to quickly respond to and control public opinion crises and reduce potential negative impacts.

[0004] Degree Centrality, Betweenness Centrality, K_Shell method, and Resistance Distance are important indicators for evaluating the influence of complex network nodes. Degree Centrality measures the dissemination ability of a node by the number of its connections, which is suitable for mining central nodes with many neighbors but ignores the overall topological structure of the network. Betweenness Centrality effectively makes up for the limitations of Degree Centrality, but it is sensitive to isolated nodes and performs poorly in small-scale networks. As a global measurement method, the K_Shell method evaluates influence by dividing nodes into different core layers. All nodes in the same core layer share the same K_Shell value, and nodes with the highest K_Shell value are generally considered the core of the network. However, this method tends to classify a large number of nodes into the same layer, resulting in a deviation in the evaluation of dissemination ability. In fact, the influence of nodes in the same layer may be significantly different. Although Resistance Distance can reflect the true distance between nodes, it has insufficient data fusion, completely relies on the topological structure and ignores node attribute information, and is sensitive to noise (such as incorrect edges), which may deviate from the true relationship.

[0005] Therefore, how to quickly and accurately mine the key nodes of social networks is a technical problem that needs to be solved urgently. Summary of the invention

[0006] In view of the above defects or improvement needs of the prior art, the present invention provides a social network key node mining method and a mining system thereof, which aims to quickly and accurately mine the key nodes of the social network.

[0007] To achieve the above object, the present invention provides a method for mining key nodes in a social network, which comprises:

[0008] Step S1: comprehensively consider multiple indicators, select p pre-selected nodes from all nodes in the social network to form a pre-selected node group S, wherein the multiple indicators include at least three of degree, betweenness, K_Shell value, and resistance distance;

[0009] Step S2: Select a pre-selected node from S that does not belong to the node group T and combine it with T to obtain the node group T'. Use the inverse power method to calculate the minimum eigenvalue of the Laplacian matrix after deleting T' from the social network as the score of the corresponding pre-selected node. Add the pre-selected node v with the highest score to T. T is an empty set at the initial moment.

[0010] Step S3: Repeat step S2 until the number of nodes in T reaches a preset scale, and take the nodes in T as key nodes.

[0011] Optionally, the inverse power method used is an inverse power method based on dynamic displacement adjustment, and the execution steps of the inverse power method based on dynamic displacement adjustment include:

[0012] Determine the network's adjacency matrix A, machine learning model M, residual threshold ∈, and maximum number of iterations K max , initial displacement, initial normalized vector;

[0013] Continue to iterate until the iteration ends and output the final eigenvalue

[0014] Each iteration process includes:

[0015] Solve the equation (Ap k I)v k+1 =u k , where p k 、u k are the displacement and normalized vector used in the kth iteration, v k+1 is the obtained iteration vector;

[0016] Calculate an approximate eigenvalue

[0017] Calculate the normalized vector u k+1 = v k+1 / ||v k+1 ||²;

[0018] Calculate the residual r k = ||Au k+1 - λ k u k+1 ||²;

[0019] Determine whether r satisfies k < ∈, if so, end the iteration; if not, determine whether k < K max , if not, end the iteration, if so, then the eigenvector component u k (1:d), p k and r k Input the machine learning model M, predict the change in displacement Δp, and update the displacement p k+1 = p k + Δp; Increment the iteration count by 1 and enter the next iteration.

[0020] Optionally, before adding the preselected node v with the highest score to T, first select the nodes that belong to S and do not belong to T from the first-order adjacent node set of v to obtain the set N(v), and determine whether N(v) is empty:

[0021] If it is empty, add v to T;

[0022] If it is not empty, select a node u from N(v), calculate the minimum eigenvalue λ1(L N-1 ) of the Laplacian matrix after deletion obtained by first deleting T' and then deleting u, and calculate the minimum eigenvalue λ2(L N-1 ) of the Laplacian matrix after deletion obtained by first deleting the node group obtained by deleting u and T and then deleting v. Use the multi-threaded method to traverse N(v). If there exists a node u that satisfies λ2(L N-1 ) > λ1(L N-1 ), then add the node u to T, otherwise, add v to T.

[0023] Optionally, if there are multiple nodes u that satisfy λ2(L N-1 ) > λ1(L N-1 ), then select the node u with the minimum resistance distance index and add it to T.

[0024] Optionally, the method further includes: after every preset number of steps of search, before performing the next step of search, randomly select a node that has not been selected And for any node g in T, compare the minimum eigenvalue of the post-removal Laplacian matrix of the social network after removing e with the minimum eigenvalue of the post-removal Laplacian matrix of the social network after removing g. If the former is larger, update T by replacing node g with node e; otherwise, do not update T and proceed to the next search.

[0025] Optionally, screen out p preselected nodes from all nodes of the social network, including:

[0026] Let the set of preselected node groups be S, and initially S = φ;

[0027] Calculate the respective index values of all nodes and sort them in descending order to obtain the sorted list of nodes under each index; select the first several nodes from each sorted list and take the union to obtain p preselected nodes and add them to the preselected node group S.

[0028] Optionally, screen out p preselected nodes from all nodes of the social network, including:

[0029] Let the set of preselected node groups be S, and initially S = φ;

[0030] Calculate the global graph properties of the graph, including graph density, average degree, average page rank, average weighted degree, average path length, and average clustering coefficient;

[0031] Calculate the evaluation scores of each index. Among them, the evaluation score of degree is obtained by weighted summation of the average degree and the average weighted degree, the evaluation score of the K_Shell value is obtained by weighted summation of the graph density, the average page rank, and the average degree, the evaluation score of betweenness is obtained by weighted summation of the average path length and the average clustering coefficient, and the evaluation score of resistance distance is obtained by weighted summation of the graph density, the average path length, and the average clustering coefficient; normalize the evaluation scores of all indexes to obtain the normalized weight of each index;

[0032] Based on the weights of each index, perform fusion calculation on all indexes to obtain the comprehensive score of each node;

[0033] Select the top p nodes with the highest comprehensive scores to form the preselected node group S.

[0034] The present invention also provides a social network key node mining system, which includes a memory and a processor. The memory stores a computer program. Wherein, when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0035] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0036] The present invention also provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the method described in any one of the above are implemented.

[0037] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the present invention mainly has the following beneficial effects:

[0038] In the present invention, first, multiple indicators are comprehensively used to pre-screen nodes within a certain range, so that the selected pre-selected nodes have relatively excellent performances in multiple indicator dimensions. Thus, the drawback that the global factors of a complex network cannot be considered when only a single indicator is used can be avoided. Therefore, the selected pre-selected node group is a relatively optimal node group. By pre-selection, the screening range is narrowed. Then, the minimum eigenvalue of the post-deletion Laplacian matrix is used as the evaluation index for re-screening. The larger the minimum eigenvalue of the post-deletion Laplacian matrix, the greater the influence of the deleted node, and the more important the node is. Therefore, further using the minimum eigenvalue of the post-deletion Laplacian matrix as the evaluation index can more accurately find important nodes in the network. At the same time, since the search range is narrowed by pre-selection, the number of nodes participating in the calculation of the minimum eigenvalue of the post-deletion Laplacian matrix is reduced, so the speed of node search is improved. On the other hand, combined with the inverse power method based on dynamic displacement adjustment, compared with the QR algorithm provided by the simulation software, it has a lower time complexity, thereby further improving the speed of node search and can better meet the requirements of the time performance of the algorithm for large-scale complex networks to a certain extent. Combined with the improved greedy algorithm, it can well overcome the drawback that the traditional greedy algorithm is prone to falling into local optimum. Generally speaking, the advantage of the present invention is that it can accurately find nodes with relatively high importance and can also ensure a relatively low algorithm complexity to a certain extent.

[0039] In some embodiments, when performing pre-selection, direct pre-screening is carried out based on the sorting of the indicator values of each indicator, so that a pre-selected set can be quickly obtained and the pre-selection efficiency can be improved.

[0040] In some embodiments, during preselection, by fusing discrete graph metrics with the aid of global graph data, it is possible to make up for the limitations of a single metric, making the importance assessment of nodes more comprehensive, avoiding unilateral biases, and also enhancing the robustness of node selection, reducing the risk of a certain metric failing, making node selection more robust. The combined comprehensive score can more reasonably reflect the role of nodes in the entire network and improve the efficiency of the selected nodes in key tasks such as information dissemination and network control. The combination of multiple metrics also achieves a balance between global and local information, and the normalized weight allocation is flexible, allowing the weights to be adjusted according to actual needs. For example, the weight of the global role metric can be increased in information dissemination tasks, or the importance of metrics related to degree values can be enhanced in local optimization tasks. This flexibility enables the method to adapt to different network analysis scenarios.

[0041] In some embodiments, when updating the node group T, directly adding the preselected node v to the node group T can improve the search speed and quickly complete the search.

[0042] In some embodiments, when updating the node group T, by comparing with adjacent nodes, a dynamic adjustment and local perturbation mechanism is introduced on the basis of the traditional greedy algorithm, solving the problem that the traditional greedy algorithm is prone to falling into local optima. By finely optimizing the neighborhood of the target node group, it can make full use of local information, dynamically evaluate the feature contributions of nodes, and optimize the structural performance of the target node group. Compared with the simple greedy algorithm, this method, while ensuring efficiency, avoids suboptimal choices caused by single decisions through backtracking and comparison, further improving the quality of the solution.

[0043] In some embodiments, by regularly comparing the selected and unselected nodes, it is possible to effectively avoid the local optimum problem of the greedy algorithm, improve the global search ability of the algorithm, contribute to balancing exploration and exploitation, and further enhance the stability and robustness of the algorithm.

[0044] In some embodiments, using the improved inverse power method described above, through machine learning-driven displacement prediction, dynamic adjustment of the displacement parameter p k is achieved, thereby optimizing the entire calculation process and avoiding the problem that the initial predicted eigenvalue of the inverse power method usually depends on the experimenter's experience for adjustment, and improper selection may lead to slow convergence or even failure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of the steps of the method for mining key nodes in a social network according to an embodiment of the present invention;

[0046] Figure 2It is a flowchart of the steps of the social network key node mining method in another embodiment of the present invention;

[0047] Figure 3(a) is the NW small-world network generated in an embodiment of the present invention;

[0048] Figure 3(b) is the network property of the NW small-world network;

[0049] Figure 3(c) is the comparison of the minimum eigenvalue λ(L N-1 ) of the deleted Laplacian matrix of the node group mined by four algorithms for the NW small-world network;

[0050] Figure 3(d) is the node importance ranking obtained by the algorithm mining for the NW small-world network;

[0051] Figure 4(a) is the lastfm_asia network generated in an embodiment of the present invention;

[0052] Figure 4(b) is the network property of the lastfm_asia network;

[0053] Figure 4(c) is the comparison of the minimum eigenvalue λ(L N-1 ) of the deleted Laplacian matrix of the node group mined by four algorithms for the lastfm_asia network;

[0054] Figure 4(d) is the node importance ranking obtained by the algorithm mining for the lastfm_asia network;

[0055] Figure 5 It is the change of the number of infections over time in the NW small-world network model;

[0056] Figure 6(a) is the facebook_combine network model in an embodiment of the present invention;

[0057] Figure 6(b) is the network property of the facebook_combine network model;

[0058] Figure 6(c) is the change of the number of infections over time in the facebook_combine network model;

[0059] Figure 7 It is the node selection situation of the lastfm_asia network under 4 different algorithms. (a) is the node selection situation of the algorithm in this paper; (b) is the node selection situation of the degree algorithm; (c) is the node selection situation of the betweenness algorithm; (d) is the node selection situation of the K_Shell algorithm. Detailed implementation manner

[0060] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] Embodiment 1

[0062] As Figure 1 shown is a flowchart of the steps of a method for mining key nodes in a social network according to an embodiment of the present invention, which includes the following steps:

[0063] Step S1: Considering multiple indicators comprehensively, p preselected nodes are screened out from all nodes in the social network to form a preselected node group S. The multiple indicators include at least three of degree, betweenness, K_Shell value, and resistance distance.

[0064] Step S2: Select a preselected node that does not belong to the node group T from S and take the union with T to obtain the node group T'. The inverse power method is used to calculate the minimum eigenvalue of the Laplacian matrix after deleting T' from the social network as the score of the corresponding preselected node. The preselected node v with the highest score is added to T, and T is an empty set at the initial moment.

[0065] Step S3: Repeat step S2 until the number of nodes in T reaches a preset scale, and the nodes in T are used as key nodes.

[0066] The following is a detailed description of the steps.

[0067] First, the effect of introducing the inverse power method in the present invention is described.

[0068] When solving problems related to calculating the minimum eigenvalue of a matrix, the matlab library function eig is usually used in past work. In the present invention, we use the inverse power method to calculate the minimum eigenvalue, which can effectively reduce the number of calculations. The following explains that the inverse power method has a lower complexity than the QR algorithm provided by the simulation software for solving the minimum eigenvalue.

[0069] The eig function uses the QR algorithm. When solving the minimum eigenvalue, it is necessary to first calculate all eigenvalues, then sort these eigenvalues, and then obtain the minimum eigenvalue. This algorithm does not consider the usage scenario (such as the scenario where only specific eigenvalues are concerned), performs a lot of unnecessary calculations, and wastes a large amount of computing resources.

[0070] The inverse power method makes the approximate value gradually converge to the smallest eigenvalue through iteration, eliminating unnecessary calculations. Moreover, the algorithm of the present invention takes into account the usage scenario and uses the inverse power method to calculate only the specific eigenvalues we need. Obviously, the computational complexity of the algorithm of the present invention has great advantages.

[0071] First, analyze the computational amount of the QR method. Let the number of iterations of the QR algorithm be T, and the matrix dimension be n. The total number of multiplications is:

[0072]

[0073] When using the inverse power method to solve for the smallest eigenvalue, let the number of iterations be t, and the approximate number of multiplications required is tn(n - 1). The approximate number of multiplication calculations required to solve for the smallest eigenvalue once by this method is:

[0074]

[0075] When T = t, for calculating the smallest eigenvalue of an n-order matrix once, the number of multiplication calculations of the inverse power method is reduced by:

[0076]

[0077] In the present invention, there are a total of p nodes in the preselected node group, and the goal is to mine a node group composed of k nodes. When mining the i-th node (i = 1, 2, 3,..., k), the matrix order for which the smallest eigenvalue needs to be calculated is n - i, and the reduced number of multiplications is p·d(t, n - i). The total reduced number of multiplication calculations for mining k nodes is:

[0078]

[0079] If considering the case of a relatively large total computational amount, let k = n - 1, that is, a target mining node group composed of n - 1 nodes is selected. At this time, the reduced number of multiplication calculations is:

[0080]

[0081] The number of reduced calculations reaches n 4 order of magnitude.

[0082] The iterative formula of the traditional inverse power method is

[0083] (A - pI)v k = u k-1

[0084] where p is a fixed displacement, and u k is the normalized eigenvector. The convergence speed depends on the approximation degree of p to the target eigenvalue, and the selection of the displacement parameter often can only be manually selected by the experimenter according to experience, which greatly affects the time complexity of calculating the smallest eigenvalue.

[0085] In one embodiment, the inverse power method is further improved to obtain an inverse power method based on dynamic displacement adjustment (machine learning-driven displacement prediction) to accelerate the iteration process and improve the calculation speed.

[0086] The execution process of the improved inverse power method includes:

[0087] Determine the adjacency matrix A of the network, the initial displacement p0, the initial vector v0, the machine learning model M, the residual threshold ∈, and the maximum number of iterations K max ;

[0088] Initialize the iteration number k = 0, the current displacement p = p0, and the current vector normalization u0 = v0 / ||v0||2;

[0089] Continuously execute the iteration until the iteration ends, and output the final eigenvalue

[0090] The process of each iteration includes:

[0091] Solve the equation (A - p k I)v k+1 = u k , where A is the adjacency matrix of the network, p k , u k are the displacement and eigenvector used in the k-th iteration respectively, and v k+1 is the obtained iteration vector;

[0092] Calculate the approximate eigenvalue

[0093] Calculate the normalized vector u k+1 = v k+1 / ||v k+1 ||2;

[0094] Calculate the residual r k = ||Au k+1 - λ k u k+1 ||2;

[0095] Judge whether r k < ∈ is satisfied. If so, end the iteration; if not, judge whether k < K max is satisfied. If not, end the iteration. If so, input the eigenvector component u k (1:d), p k and r k into the machine learning model M to predict the change in displacement Δp, and update the displacement p k+1 = p k + Δp; increment the iteration number by 1 and enter the next iteration.

[0096] More specifically, it can be programmed as follows:

[0097] Input: matrix A, initial displacement p0, initial vector v0, machine learning model M, residual threshold ∈, maximum number of iterations K max ;

[0098] Output: eigenvalue λ, eigenvector u;

[0099] 1. Initialization: iteration number k = 0, current displacement p = p0, current vector normalization u0 = v0 / ||v0||2;

[0100] 2. while k < K max ;

[0101] 3. Solve the equation (A - p k I)v k+1 - u k (Perform LU decomposition on (A - pI) to accelerate the solution and obtain the new iterative vector v k+1 );

[0102] 4. Calculate the approximate eigenvalue (Obtain the estimated value of the eigenvalue λ based on the iterative vector k );

[0103] 5. Vector normalization u k+1 = v k+1 / ||v k+1 ||2 (Normalize the new iterative vector to obtain the vector u for the next loop k+1 );

[0104] 6. Calculate the residual r k = ||Au k+1 - λ k u k+1 ||2 (Calculate the residual r k , compare it with the error threshold ∈ to determine convergence);

[0105] 7. if r k < ∈ break;

[0106] 8. Extract the eigenvector components u k (1:d) (the first d components of u k ), residual r k ;

[0107] 9. Predict Δp = M(p k , u k (1:d), r k )(Obtain the predicted value of the displacement change);

[0108] 10. Update displacement p k+1 = p k + Δp;

[0109] 11. Iteration count k = k + 1;

[0110] 12. end while;

[0111] 13. return andu = v k+1 (Output the finally obtained eigenvalue and vector).

[0112] In this embodiment, the improved inverse power method described above is adopted, and through machine learning-driven displacement prediction, the dynamic adjustment of the displacement parameter p k is realized, thereby optimizing the entire calculation process and avoiding the problem that the initial predicted eigenvalue of the inverse power method usually depends on the experimenter's experience for adjustment, and improper selection may lead to slow convergence or even failure.

[0113] Specifically, the machine learning model can adopt a random forest regression model, which effectively balances the calculation efficiency and prediction accuracy by integrating the prediction results of multiple decision trees and is suitable for numerical optimization scenarios with high dimensions and small samples.

[0114] The purpose of step S1 is to pre-screen nodes that meet multi-index requirements from the social network. The multi-index includes at least three of degree, betweenness, K_Shell value, and resistance distance. Degree, betweenness, K_Shell value, and resistance distance are all indicators for evaluating node importance, but they each have disadvantages. Evaluating nodes based on any one indicator will result in biases. In the present invention, first, multiple indicators are comprehensively used to pre-screen nodes within a certain range, so that the selected pre-selected nodes have good performance in multiple index dimensions. Thus, the drawback of not being able to consider the global factors of the complex network when using a single indicator can be avoided. Since the pre-selected nodes have good performance in multiple index dimensions, the probability that they belong to key nodes is also relatively high. They are put into the pre-selected node group to narrow the scope of further screening. Subsequently, as long as the pre-selected node group is further screened, nodes with higher importance can be quickly found and have a higher mining accuracy.

[0115] In one embodiment, p pre-selected nodes can be screened from all nodes of the social network through the following steps S101 to S102:

[0116] Step S101: Let the set of the pre-selected node group be S, and initially S = φ;

[0117] Step S102: Calculate the respective index values of all nodes and sort them in descending order to obtain the sorted list of nodes under each index; select the top several nodes from each sorted list, take the union, and add them to the preselected node group S.

[0118] Further, selecting the top several nodes from each sorted list, taking the union, and adding them to the preselected node group S includes:

[0119] Select the top several nodes from each sorted list and take the union to obtain the initial union set; use the inverse power method to calculate the minimum eigenvalue of the Laplacian matrix after deleting each node in the initial union set itself, and select the p nodes that make the minimum eigenvalue the largest and add them to the preselected node group S.

[0120] In the above embodiments, pre-screening is directly based on the sorting of the index values of each index, and the preselected set can be quickly obtained, improving the preselection efficiency.

[0121] Specifically, the network can be represented by a Laplacian matrix. The sum of each row of it is 0, and the diagonal elements are the degrees of the nodes. If we are concerned about the importance of a node group composed of certain points, then the corresponding rows and columns of the Laplacian matrix can be deleted, and the remaining principal submatrix is called the matrix after deletion. The minimum eigenvalue of the Laplacian matrix after deletion can also be used to evaluate the importance of nodes. The larger the minimum eigenvalue of the Laplacian matrix after deletion, the more important the deleted nodes are.

[0122] In another embodiment, the following steps S111 to S115 can also be used to screen p preselected nodes from all nodes in the social network.

[0123] Step S111: Let the set of preselected node groups be S, and initially S = φ

[0124] Step S112: Calculate the global graph properties of the graph, including graph density, average degree, average PageRank, average weighted degree, average path length, and average clustering coefficient.

[0125] First, perform global graph information extraction.

[0126] Specifically, calculate the global graph properties of the graph, including graph density density, average degree avg_degree, average PageRank, average weighted degree avg_weighted_degree, average path length avg_path_length, and average clustering coefficient avg_clustering_coefficient.

[0127] Step S113: Calculate the evaluation scores of each indicator. Among them, the evaluation score of degree is obtained by summing the average degree and the average weighted degree; the evaluation score of the K_Shell value is obtained by summing the graph density, the average page rank, and the average degree; the evaluation score of betweenness is obtained by performing a weighted sum of the average path length and the average clustering coefficient; the evaluation score of resistance distance is obtained by summing the graph density, the average path length, and the average clustering coefficient; normalize the evaluation scores of all indicators to obtain the normalized weight of each indicator.

[0128] Specifically, the evaluation score f degree of degree, the evaluation score f K_Shell of the K_Shell value, the evaluation score f BC of betweenness, and the evaluation score f Ris_Dis of resistance distance are calculated as follows:

[0129] f degree = a1·avg_degree + a2·avg_weighted_degree;

[0130] f K_Shell = b1·density + b2·PageRank + b3·avg_degree;

[0131] f BC = c1·avg_path_length + c2·avg_clustering_coefficient;

[0132] f Res_dis = d1·density + d2·avg_path_length + d3·avg_clustering_coefficient;

[0133]

[0134] In the formula, a1, a2, b1, b2, b3, c1, c2, d1, d2, and d3 are all adjustable parameters used to adjust the importance of different indicators in the calculation.

[0135] Then, normalize all the evaluation scores to obtain the normalized weight of each indicator:

[0136]

[0137] In the formula, f i is the evaluation score of the i-th indicator, and w i is the normalized weight of the i-th indicator.

[0138] Step S114: Based on the weights of each indicator, all indicators are fused and calculated to obtain the comprehensive score of each node.

[0139] Specifically, let the comprehensive score of node v be F(v), and the fusion method can be expressed as:

[0140]

[0141] In the formula, feature i (v) represents the value of the i-th indicator of node v, w i is the normalized weight of the i-th indicator, represents the combination method of feature i and w i , and ’Θ represents the operation on the results of all indicators again. The result of

[0142] For example, the fusion method can be specifically expressed as:

[0143]

[0144] In the formula, degree(v), K_Shell(v), Res_Dis(v), and BC(v) represent the degree, K_Shell, resistance distance, and betweenness of node v respectively, and w degree , w K_Shell , w Res_Dis , w BC represent the weights of degree, K_Shell, resistance distance, and betweenness respectively. In this formula, the weight can be used as the exponent of the indicator or as the coefficient of the indicator.

[0145] Step S115: Select the top p nodes with the highest comprehensive scores to form the preselected node group S.

[0146] In one embodiment, after step S114, sorting can be directly performed, and then the top p nodes with the highest comprehensive scores are selected to form the preselected node group S, which can speed up the preselection speed and ensure the search efficiency.

[0147] In the above embodiments, by leveraging global graph data for discrete graph index fusion, it is possible to make up for the limitations of a single index, enabling a more comprehensive evaluation of the importance of nodes, avoiding unilateral biases, and enhancing the robustness of node selection. The risk of a certain index failing is reduced, making node selection more stable. The combined comprehensive score can more reasonably reflect the role of nodes in the entire network, improving the efficiency of the selected nodes in key tasks such as information dissemination and network control. The combination of multiple indices also achieves a balance between global and local information, and the normalized weight allocation is flexible, allowing the weights to be adjusted according to actual needs. For example, the weight of the global role index can be increased in information dissemination tasks, or the importance of degree value-related indices can be enhanced in local optimization tasks. This flexibility enables the method to adapt to different network analysis scenarios.

[0148] After determining the preselected node group S, steps S2 to S3 are executed.

[0149] The purpose of steps S2 to S3 is to further screen the preselected node group S to obtain a preset number of the most important nodes.

[0150] Step S2 is repeatedly executed. Each execution of step S2 is a search.

[0151] Specifically, the node group T is initialized to be empty. Traverse the nodes in the preselected node group S that have not been selected into the node group T. For each selected node, calculate the union with the node group T to obtain the node group T'. Use the inverse power method to calculate the smallest eigenvalue λ(L N-1 ) of the Laplacian matrix after deleting the node group T'. The larger the smallest eigenvalue of the Laplacian matrix after deletion, the greater the influence of the deleted node, and the more important the node. Select the preselected node v that maximizes the smallest eigenvalue, and update the node group T based on the preselected node v.

[0152] Repeat the search, expanding the scale of the nodes in the node group T until the number of nodes in the node group T reaches the preset number requirement, and output the nodes in the node group T as the mined key nodes.

[0153] The present invention uses the minimum eigenvalue of the Laplacian matrix after deletion as the evaluation index for re-screening. The larger the minimum eigenvalue of the Laplacian matrix after deletion, the greater the influence of the deleted node, and the more important the node. Using the minimum eigenvalue of the Laplacian matrix after deletion as the evaluation index can accurately find important nodes in the network. At the same time, in order to improve the speed of obtaining important nodes, the present invention first pre-selects relatively important nodes through multiple indicators to form a pre-selected node group, thereby narrowing the search range, and then successively searches based on the idea of greedy search and calculates the minimum eigenvalue of the Laplacian matrix after deletion using the inverse power method. On the one hand, since the search range is narrowed at the beginning, the number of nodes participating in the calculation of the minimum eigenvalue of the Laplacian matrix after deletion is reduced, thereby improving the speed of node search. On the other hand, combined with the inverse power method based on dynamic displacement adjustment, compared with the QR algorithm built in the simulation software, it has a lower time complexity (because the QR algorithm calculates all eigenvalues and then finds the minimum eigenvalue through sorting, while the inverse power method based on dynamic displacement adjustment directly iteratively finds the minimum eigenvalue, and the number of iterations is greatly reduced compared with the traditional inverse power method), thereby improving the speed of node search and being able to better meet the requirements of large-scale complex networks for the time performance of the algorithm to a certain extent. Therefore, the advantage of the algorithm of the present invention is that it can accurately find nodes with higher importance and at the same time ensure a relatively low algorithm complexity to a certain extent.

[0154] In one embodiment, the pre-selected node v can be directly added to the node group T and then directly enter the next search.

[0155] In this embodiment, directly adding the pre-selected node v to the node group T can improve the search speed and quickly complete the search.

[0156] In another embodiment, as Figure 2 shown, before adding the pre-selected node v with the highest score to T, first select the nodes that belong to S and do not belong to T from the first-order adjacent node set of v to obtain the set N(v), and judge whether N(v) is empty: if it is empty, add v to T; if it is not empty, select a node i from N(v), and calculate the minimum eigenvalue λ1(L N-1 ) of the Laplacian matrix after deletion obtained by first deleting T' and then deleting u, and calculate the minimum eigenvalue λ2(L N-1 ) of the Laplacian matrix after deletion obtained by first deleting the node group obtained by deleting u and T and then deleting v. Traverse N(v), if there exists a node u that satisfies λ2(L N-1 ) > λ1(L N-1 ), add the node u to T, otherwise, add v to T.

[0157] In this embodiment, by comparing with adjacent nodes, a dynamic adjustment and local perturbation mechanism is introduced on the basis of the traditional greedy algorithm, solving the problem that the traditional greedy algorithm is prone to falling into local optimum. By finely optimizing the neighborhood of the target node group, it can make full use of local information, dynamically evaluate the feature contributions of nodes, and optimize the structural performance of the target node group. Compared with the simple greedy algorithm, while ensuring efficiency, this method avoids suboptimal choices caused by single decisions through backtracking and comparison, further improving the quality of the solution.

[0158] Further, if there are multiple nodes u that satisfy λ2(L N-1 )>λ1(L N-1 ), then the node u with the smallest resistance distance index is selected from them and added to T.

[0159] The following explains the resistance distance.

[0160] Considering the given network graph G=(V, E) and its Laplacian matrix L N , if H is a generalized inverse of L N , then the resistance distance between nodes i and j is defined as follows:

[0161] r(i, j)e ij T He ij =h ii +h jj -h ij -h ji

[0162] where e ij is an N×1 column vector, with the element in the i-th row being 1, the element in the j-th row being -1, and the elements in other rows being 0. When i = j, r(i, j) is 0. Different from the physical distance between nodes, the resistance distance can be understood as the equivalent resistance between nodes. Where h ij represents the i-th row and j-th column of the H matrix.

[0163] And the resistance distance σ t of a certain node t in the network is equal to the sum of the resistance distances from this node t to other nodes:

[0164]

[0165] For the above network and Laplacian matrix, if node t is the only controlled node, then:

[0166]

[0167] where, λ N (L N )≥...≥λ1(LN ), namely L N All eigenvalues are arranged from largest to smallest. It can be seen from this formula that when the resistance distance σ t is smaller, the right part of the formula is larger, the upper bound of the smallest eigenvalue is larger, and correspondingly, λ1(L N-1 ) may also be larger.

[0168] Therefore, the resistance distance of nodes can also be used to measure the importance of nodes. Thus, if there are multiple nodes u that satisfy λ2(L N-1 ) > λ1(L N-1 ), then select the node u with the smallest resistance distance index from them and add it to T. The smaller the resistance distance index, the larger the smallest eigenvalue of the Laplacian matrix after its deletion may be, indicating that the node is more likely to be important. In this way, more important nodes can be mined.

[0169] In one embodiment, after every N steps of search, before performing the next search, randomly select a node that has not been selected before and any node g from T, compare the smallest eigenvalue of the Laplacian matrix after deleting e from the social network and the smallest eigenvalue of the Laplacian matrix after deleting g from the social network. If the former is larger, then update T by replacing node g with node e; otherwise, do not update T and proceed to the next search.

[0170] In this embodiment, by regularly comparing the selected and unselected nodes, the local optimal problem of the greedy algorithm can be effectively avoided, the global search ability of the algorithm is improved, which helps to balance exploration and exploitation, and can further enhance the stability and robustness of the algorithm.

[0171] The effects of the present invention are described below through specific embodiments.

[0172] (1) Experiment on the smallest eigenvalue of the Laplacian matrix after deletion

[0173] The size of the smallest eigenvalue of the Laplacian matrix after deletion of a node group is related to the importance of the node group. The larger the smallest eigenvalue of the Laplacian matrix after deletion, the greater the influence of the node group on the entire network, and the more important the node group is in the complex network.

[0174] Therefore, taking the smallest eigenvalue of the Laplacian matrix after deletion as an index, we compare the effects of four algorithms, namely the contrast algorithm, the betweenness algorithm, the K_Shell algorithm, and the algorithm of the present invention, in screening important node groups. If the smallest eigenvalue of the Laplacian matrix after deletion corresponding to the screened important node group is larger, it indicates that the node group mined by the algorithm has a better effect on the network containment control, and the algorithm effect is better.

[0175] Network 1: Consider a NW small-world network model with the number of nodes N = 1000, where each node is connected to its nearest L N = 4 nodes, and the edges are changed with a random probability p = 0.1. The generated NW small-world network is shown in Fig. 3(a), and the network properties are shown in Fig. 3(b). Among them, the nodes with larger degrees are larger in size. Use the degree algorithm, betweenness algorithm, K_Shell algorithm, and the algorithm of this paper to mine important node groups. Increase the scale k of the mined important node groups in turn, and calculate the minimum eigenvalue λ(L N-1 ) of the post-deletion Laplacian matrix after deleting the node groups mined by the above four algorithms, as shown in Fig. 3(c). The sorting of the node importance obtained by each algorithm is shown in Table 3(d).

[0176] Network 2: A real network lastfm_asia with a size of N = 7624. The network node graph is shown in Fig. 4(a), and the network properties are shown in Fig. 4(b). Use the degree algorithm, betweenness algorithm, K_Shell algorithm, and the algorithm of this paper to mine important node groups. Increase the scale k of the mined important node groups in turn. As shown in Fig. 4(c), the comparison of the minimum eigenvalues λ(L N-1 ) of the post-deletion Laplacian matrix after deleting the node groups mined by the above four algorithms is shown. The sorting of the node importance obtained by each algorithm is shown in Table 4(d).

[0177] It can be seen from Fig. 3(c) and Fig. 4(c) that the minimum eigenvalue λ(L N-1 ) of the post-deletion Laplacian matrix corresponding to the node group screened by the algorithm of the present invention always leads the other three algorithms. It proves that the node group screened by it has significant advantages in complex network control.

[0178] (2) SIR model simulation analysis

[0179] In the SIR model, use the screened node group as the initial infected nodes for simulation. If the number of infections rises faster and the extreme value is larger at the same time step, it means that the algorithm has a better effect in mining nodes with stronger propagation ability.

[0180] The SIR model is often used to describe a certain propagation state, and the infected individuals can acquire immunity after being cured. The model defines three types of individuals in different states, namely the susceptible state S, the infected state I, and the removed state R. There are the following transformation relationships among the individuals in the three states: individuals in the susceptible state have a certain probability of being infected by their neighbor individuals (individuals with frequent interactions); infected individuals have a probability of being cured and becoming individuals in the removed state. When certain nodes are selected as the sources of infection, the rising speed of the number of infected individuals and the extreme value of the number of infected individuals can reflect the importance of these nodes in network propagation. Therefore, using the selected node group as the initial infected nodes for simulation, if the number of infections rises faster and the extreme value is larger within the same number of time steps, it indicates that the algorithm has a better effect in mining nodes with stronger propagation capabilities.

[0181] Model 1: Using the NW small-world network model with the same network information above, taking the important node groups in Figure 3(d) as the initial infected node groups respectively, setting the infection rate β = 0.01 and the cure rate γ = 0.005 for infection simulation (the number of infections is the average value of ten experiments), the change of the number of infections over time is as Figure 5 shown.

[0182] Model 2: Considering another form of the facebook_combine real social network with the number of nodes N = 1519, whose network structure is shown in Figure 6(a) and network properties are shown in Figure 6(b). The infection rate β and the cure rate γ are set as above. Using four algorithms to mine important node groups, and then taking the mined important node groups as the initial infected individuals for infection simulation experiments, the results are shown in Figure 6(c).

[0183] From Figure 5 and Figure 6(c), it can be seen that the algorithm of the present invention can more quickly achieve the spread of infection in the initial stage of infection, which is reflected in the image as: the curve has a larger extreme value and a larger growth rate, indicating that the selected node group has a more important position in the network.

[0184] Figure 7 Shows the important node groups selected by four algorithms in the lastfm_asia network, where the red dots are the selected node groups. (a) shows the node selection situation of the algorithm in this paper; (b) shows the node selection situation of the degree algorithm; (c) shows the node selection situation of the betweenness algorithm; (d) shows the node selection situation of the K_Shell algorithm. The algorithm of the present invention can more comprehensively mine potential key node groups by comprehensively considering the degree, betweenness centrality, K-shell value, and resistance distance of nodes, taking into account both global and local information, which makes it perform excellently in terms of propagation efficiency and network robustness.

[0185] Generally speaking, the method proposed by the present invention can accurately calculate the optimal important node group. Such nodes usually have a relatively large degree, are roughly in the core position in the network, are located in a relatively central shell layer, and are relatively dispersed in position, which is an optimal important node group.

[0186] Embodiment 2

[0187] The present invention also relates to a social network key node mining system, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0188] The system can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The so-called processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.

[0189] Embodiment 3

[0190] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0191] Specifically, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0192] Embodiment 4

[0193] An embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method in the above embodiment of the present invention.

[0194] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. It should be noted that the "in an embodiment of the present invention", "for example", "again, for example", etc. in the present invention are intended to illustrate the present invention, rather than to limit the present invention.

[0195] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the scope of the patent application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for mining key nodes in a social network, characterized in that: include: Step S1: comprehensively consider multiple indicators, select p pre-selected nodes from all nodes in the social network to form a pre-selected node group S, wherein the multiple indicators include at least three of degree, betweenness, K_Shell value, and resistance distance; Step S2: Select a pre-selected node from S that does not belong to the node group T and combine it with T to obtain the node group T'. Use the inverse power method to calculate the minimum eigenvalue of the Laplacian matrix after deleting T' from the social network as the score of the corresponding pre-selected node. Add the pre-selected node v with the highest score to T. T is an empty set at the initial moment. Step S3: Repeat step S2 until the number of nodes in T reaches a preset scale, and take the nodes in T as key nodes.

2. The method for mining key nodes of a social network as claimed in claim 1, characterized in that: The inverse power method used is an inverse power method based on dynamic displacement adjustment. The execution steps of the inverse power method based on dynamic displacement adjustment include: Determine the network's adjacency matrix A, machine learning model M, residual threshold ∈, and maximum number of iterations K max , initial displacement, initial normalized vector; Continue to iterate until the iteration ends and output the final eigenvalue Each iteration process includes: Solve the equation (Ap k I)v k+1 =u k , where p k 、u k are the displacement and normalized vector used in the kth iteration, v k+1 is the obtained iteration vector; Calculate an approximate eigenvalue Calculate the normalized vector u k+1 =v k+1 / ||v k+1 ||2; Calculate the residual r k =||Au k+1 -λ k u k+1 ||2; Determine whether r is satisfied k <∈, if yes, then end the iteration; if not, then determine whether k<K max If not, the iteration ends. If so, the eigenvector component u k (1:d), p k and r k Input the machine learning model M, predict the change in displacement Δp, and update the displacement p k+1 =p k +Δp; the number of iterations increases by 1 and enters the next iteration.

3. The method for mining key nodes of a social network as claimed in claim 1, characterized in that: Before adding the pre-selected node v with the highest score to T, first select the nodes that belong to S and not to T from the first-order adjacent node set of v to obtain the set N(v), and determine whether N(v) is empty: If it is empty, add v to T; If it is not empty, select a node u from N(v) and calculate the minimum eigenvalue λ1(L N-1 ) and calculate the minimum eigenvalue λ2(L of the Laplacian matrix obtained by first deleting u and T and then deleting v N-1 ), traverse N(v), if there exists a condition that satisfies λ2(L N-1 )>λ1(L N-1 ), then add node u to T, otherwise, add v to T.

4. The method for mining key nodes of a social network as claimed in claim 3, characterized in that: If there are multiple N-1 )>λ1(L N-1 ) of the node u, then select the node u with the smallest resistance distance index and add it to T.

5. The method for mining key nodes of a social network as claimed in claim 1, characterized in that: The method further comprises: after each preset step of searching, before executing the next step of searching, randomly selecting a node that has not been selected And from any node g in T, compare the minimum eigenvalue of the deleted Laplacian matrix after deleting e from the social network with the minimum eigenvalue of the deleted Laplacian matrix after deleting g from the social network. If the former is larger, update T with node e instead of node g. Otherwise, do not update T and proceed to the next step of search.

6. The method for mining key nodes in a social network as claimed in claim 1, characterized in that: Filter out p pre-selected nodes from all nodes in the social network, including: Let the set of pre-selected node groups be S, and the initial S = φ; Calculate the index values ​​of all nodes respectively and sort them in descending order to obtain a sorted list of nodes under each index; select the first several nodes from each sorted list respectively, obtain p pre-selected nodes and add them to the pre-selected node group S.

7. The method for mining key nodes in a social network as claimed in claim 1, characterized in that: Filter out p pre-selected nodes from all nodes in the social network, including: Let the set of pre-selected node groups be S, and the initial S = φ; Compute global graph properties of the graph, including graph density, average degree, average page rank, average weighted degree, average path length, and average clustering coefficient; Calculate the evaluation score of each indicator, wherein the evaluation score of degree is obtained by weighted summing the average degree and the average weighted degree, the evaluation score of K_Shell value is obtained by weighted summing the graph density, the average webpage ranking and the average degree, the evaluation score of betweenness is obtained by weighted summing the average path length and the average clustering coefficient, and the evaluation score of resistance distance is obtained by weighted summing the graph density, the average path length and the average clustering coefficient; normalize the evaluation scores of all indicators to obtain the normalized weight of each indicator; All indicators are integrated and calculated based on the weight of each indicator to obtain the comprehensive score of each node; The first p nodes with the highest comprehensive scores are selected to form the pre-selected node group S.

8. A social network key node mining system, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.