A Method for Measuring the Importance of Social Network Nodes Based on Multi-Feature Fusion

Through the improved D-S evidence theory, the centrality, transitiveness and reputation indicators of social network nodes are integrated, and the problem of low accuracy in handling evidence conflicts and vague situations is solved, and the high accuracy of the importance measurement for large social network nodes is achieved.

CN114913027BActive Publication Date: 2025-06-17ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210488358.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2025-06-17
Estimated Expiration
2042-05-06

AI Technical Summary

Technical Problem

The existing social network node importance measurement method is not very accurate when dealing with evidence conflicts and vague situations, making it difficult to apply to real large social networks.

Method used

The improved D-S evidence theory is adopted, and the evidence is fusion using combination rules to obtain the importance of nodes by integrating indicators such as centrality, transitiveness and reputation of nodes.

Benefits of technology

It effectively alleviates the impact of data conflicts and blurring on node importance measurement in real social networks, improves measurement accuracy, and is suitable for large social networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913027B_ABST
    Figure CN114913027B_ABST
Patent Text Reader

Abstract

A method for measuring the importance of social network nodes based on multi-feature fusion uses centrality indicators, transitivity indicators, and prestige indicators as attribute features for measuring the importance of nodes; secondly, it fuses the three features. In this process, a probability assignment generation method based on the interval number model is used to transform the feature values into basic probability assignment functions, and then the improved evidence theory is used to fuse multiple BPAs; finally, the probability of node importance is obtained based on the fused indicators, and the nodes are sorted. When measuring the importance of nodes in a real large-scale social network, this method has high measurement accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network analysis, and specifically relates to a method for measuring the importance of social network nodes based on multi-feature fusion. Background Art

[0002] The method for measuring the importance of social network nodes based on multi-feature fusion refers to the process of converting social network user feature data (including centrality, transitivity, prestige, etc.) into evidence, and then using the D-S evidence theory to fuse them to obtain a metric index for judging the importance of nodes, so as to measure the node importance probability. The specific process is as Figure 1 shown.

[0003] Node importance measurement is mainly used to find nodes with great propagation influence in the field of social networks. Such nodes can affect or even determine the structure and function of the entire network, and can quickly affect most other nodes on the network. Timely identifying important nodes has practical significance for effectively controlling social network resources, guiding the development direction of network public opinion, and controlling the spread of rumors.

[0004] In a real social network, due to the influence of network coupling information and transmission mechanisms, the centrality, transitivity, and prestige of nodes often present an uncertain relationship, and the information among the three may cancel each other out. And after converting one of the attribute features into evidence, there may be conflicts between any two of them, or it may be difficult to determine the importance degree of nodes based on one of them. Previous studies only used the traditional D-S evidence theory to solve the problem of mutual cancellation during the fusion of attribute features converted into evidence, but did not consider comprehensively the phenomenon of mutual conflict between evidences or highly fuzzy data within evidences, thus affecting the accuracy of node importance measurement, resulting in the measurement method being inapplicable to real large-scale social networks. Therefore, there is room for improvement in the accuracy of social network node importance measurement. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for measuring the importance of social network nodes based on multi-feature fusion. In this method, the D-S evidence theory that can effectively handle evidence conflicts and fuzzy situations is used to fuse indicators such as the centrality, transitivity, and prestige of nodes, and the nodes are sorted according to the fusion results. This method can effectively solve the problem that data conflicts and fuzziness in a real network affect the accuracy of node importance measurement.

[0006] To achieve the above technical purpose, the technical solution adopted is: A method for measuring the importance of social network nodes based on multi-feature fusion, comprising the following steps:

[0007] Step 1, calculate the centrality index, transitivity index, and prestige index of nodes respectively;

[0008] Step 2, represent as the probability that node v is important or unimportant under the criterion based on the centrality index; m i ; m F,I (i), m F,NI (i) is represented as the probability that node v is important or unimportant under the criterion based on the transitivity index; i ; is represented as the probability that node v is important or unimportant under the criterion based on the prestige index; i ;

[0009] Step 3, fuse m m F,I (i), m F,NI (i), using the combination rule for evidence fusion;

[0010] The rule for evidence fusion in Step 3 is to fuse m F,I (i), into the importance probability of node vi, and fuse m F,NI (i), into the unimportance probability of node vi. The importance or unimportance probability of node vi is obtained by subtracting the importance probability and the unimportance probability of node vi from 1;

[0011] It further includes Step 4, assign the importance or unimportance probability of node vi to the importance probability of node v i to obtain the MFF(i) value of each node, and rank the nodes in the social network in descending order according to the MFF(i) value of the nodes. The nodes ranked higher are more important.

[0012] Furthermore, the centrality index is obtained by dividing the out-degree of the node by the total number of nodes minus 1.

[0013] Furthermore, the transitivity index is represented by the linear negative correlation function of the local clustering coefficient of the node.

[0014] Furthermore, the prestige index is obtained by dividing the in-degree of the node by the total number of nodes minus 1.

[0015] Furthermore, is constructed based on the centrality index, the maximum value of the centrality index, the minimum value of the centrality index, and the modification parameter λ using the probability assignment function. m F,I (i), m F,NI (i) is constructed based on the transitivity index, the maximum value of the transitivity index, the minimum value of the transitivity index, and the modification parameter δ using the probability assignment function; It is constructed based on a probability assignment function using a prestige index, a maximum value of the prestige index, a minimum value of the prestige index, and a correction parameter μ.

[0016] Furthermore, the MFF(i) value of each node is obtained according to the following formula

[0017]

[0018] where m I (i) represents the importance probability of node v i and represents the probability that node v i is important or unimportant.

[0019] The beneficial effects of the present invention are as follows: It alleviates the influence of network coupling information and information transmission mechanisms on the measurement of node importance in real social networks, and solves the problem that there may be conflicts between pairs or it is difficult to determine the importance degree of nodes based on one of them after the attribute features are converted into evidence. When measuring the importance of nodes in real large-scale social networks, this method has high measurement accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a framework diagram of a multi-feature fusion method in the prior art;

[0021] Figure 2 is a framework diagram of the fusion method of the present invention;

[0022] Figure 3 is a diagram of the label table of the top 20 important nodes obtained by four methods in the specific experiment of the present invention;

[0023] Figure 4 is a diagram of the experimental results of the robustness and vulnerability of the present invention;

[0024] Figure 5 is a comparison diagram of the experimental results of the SIR propagation characteristics of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] As Figure 2 shown, a framework diagram of a method for measuring the importance of nodes in a social network based on multi-feature fusion is as Figure 2 shown. Since in the field of social network analysis, centrality, transitivity, and prestige are usually used to describe the topological characteristics of the network and the interaction behaviors between nodes, and the three constitute the influencing factors of node importance in the social network. Therefore, centrality indicators, transitivity indicators, and prestige indicators are used as attribute features for measuring node importance; secondly, the three features are fused. In this process, the original feature data (original evidence) is preprocessed first, that is, a probability assignment generation method based on the interval number model is used [1]Convert the eigenvalue into a BPA (basic probability assignment function), then assign weights to the evidence based on the evidence conflict and fuzziness to correct the original BPA. After that, use the Dempster combination rule to fuse the corrected evidence N - 1 times to obtain the fusion index (N is the number of pieces of evidence); finally, obtain the probabilities that the nodes are important or unimportant based on the fused index, and rank the nodes. The specific steps are as follows (1)-(5):

[0026] (1) Calculate the relevant indicators

[0027] Calculate the centrality index of the node respectively Transitivity index F i and prestige index

[0028]

[0029]

[0030] Among them, represents the centrality of node v i , represents the out-degree of node v i , N represents the total number of nodes. Γ(i) represents the set of neighbor nodes of node v i , a ij represents the connection status between nodes. The value of 0 indicates no connection between nodes, and the value of 1 indicates a connection with an out-degree pointing between nodes.

[0031]

[0032]

[0033]

[0034] Among them, C M,i represents the local clustering coefficient of the node, 2N jk represents the number of edges between node v i and its neighbor nodes, d i represents the degree of the node, represents node v i and the degree of mutual pointing between neighbor nodes. b ij represents the connection status between nodes. The value of 0 indicates no connection between nodes, and the value of 1 indicates a connection between nodes. F is a linear negative correlation function of C M,i and is usually used to represent the transitivity index.

[0035]

[0036]

[0037] Among them, represents the prestige index of node v i , represents the in-degree of node v i , and c ij represents the connection status between nodes. The value of 0 indicates no connection between nodes, and the value of 1 indicates a connection with an in-degree between nodes.

[0038] And respectively count the maximum and minimum values of each type of index, denoted as F M , F m , Under normal circumstances, F M ≠F m ,

[0039] (2) Determine the identification framework of nodes

[0040] In the results of judging the importance of social network nodes, there are two situations, namely important or unimportant. Therefore, the identification framework of nodes φ = {I, NI}, where I represents that the node is an important node, and NI represents that the node is an unimportant node. The two are mutually exclusive and φ is a finite complete set.

[0041] (3) Construct the basic probability assignment function (BPA function)

[0042] Express the function as the probability that node v i is important or unimportant under the standard based on the centrality index; the function m F,I (i), m F,NI (i) represents the probability that node v i is important or unimportant under the standard based on the transitivity index; the function represents the probability that node v i is important or unimportant under the standard based on the prestige index. The function m F,I (i), The larger the value of, the more important the node is under this standard. Without specific conditions, usually constrain this BPA function to follow a uniform distribution, and the expressions of each BPA function are as follows:

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049] Among them, λ, δ, and μ are correction parameters, and their values have no influence on node sorting. λ, δ, μ ∈ (0, 1).

[0050] (4) Perform multi - evidence fusion

[0051] Using the improved combination rule of evidence theory, in the evidence fusion stage, the BPA functions under centrality, transitivity, and prestige are used with the new combination rule for evidence fusion, and three new BPA functions can be obtained, denoted as m I (i), m NI (i), where m I (i) represents the importance probability of node v i , m NI (i) represents the unimportance probability of node v i , represents the probability that node v i is important or unimportant. m I (i), m NI (i), The calculation formulas are as follows:

[0052]

[0053]

[0054]

[0055] (5) Obtain the MFF of each node

[0056] represents the probability that node v i is important or unimportant. In the case of no prior knowledge, according to Bayesian theory, this probability can be assigned to m I (i), m NI (i) to further correct the probability that the node is important or unimportant. The correction method is as follows:

[0057]

[0058]

[0059] The calculated MFF(i) is the basis for finally measuring the importance of nodes. Nodes in the network are sorted in descending order according to their MFF(i) values. The nodes with higher rankings are more important. In the specific application of social network rumor control, usually, some important nodes are selected according to the size of the network structure as rumor-refuting nodes to spread the truth and combat rumors. Using important nodes with great propagation influence as rumor-refuting nodes can enhance the infection strength and propagation range of true information, so as to achieve the purpose of reducing rumor-believing users and curbing rumor propagation.

[0060] In the specific implementation stage, real social network datasets (Higgs Twitter Dataset, Advogato, Wiki-vote) were selected to verify the effectiveness of this method. Among them, the Higgs Twitter Dataset was obtained from the Twitter platform around a specific event and contains 4 directed relationships: following, retweeting, mentioning, and replying, that is, 4 directed networks can be constructed (higgs-retweet_network, higgs-mention_network, higgs-reply_network, higgs-social_network); Advogato and Wiki-vote are public datasets, respectively, from which a user trust relationship network and a Wikipedia voting network can be constructed. The basic statistical information of the datasets is shown in Table 1.

[0061] Table 1 Basic statistical information of the experimental datasets

[0062]

[0063] Randomly select a node in the Advogato user relationship network and measure the importance of this node. Suppose the selected node is the node with ID number 57. The values of each index calculated for this node are shown in Table 2:

[0064] Table 2 Nodes V in Advogato_network 57 Numerical values of each index

[0065]

[0066] In order to scientifically and comprehensively evaluate the proposed method, a comprehensive evaluation of MFF based on network robustness and vulnerability and SIR propagation characteristics was carried out on the above real network datasets. The comparison methods selected are: multiplicity centrality (MEC) that only focuses on network centrality; multiple eigenvector centrality (MEVC) that focuses on network centrality, transitivity, and prestige but is based on the traditional d-s evidence theory.

[0067] (1) Multiplicity centrality (MEC): MEC = K out⊕F;

[0068] Among them, K out represents the out-degree centrality of a node, F represents the linear negative correlation function of the local clustering coefficient of a node, and ⊕ represents fusion using the original evidence theory combination rule.

[0069] (2) Multiple Eigenvector Centrality (MEVC): K out ⊕F⊕K in ;

[0070] Among them, K in represents the in-degree centrality of a node, and K out , F, and K in are fused using the original evidence theory combination rule.

[0071] (3) The method of the present invention: Multi-Feature Fusion Method (MFF): K out ⊕F⊕K in

[0072] Among them, K out , F, and K in are fused using the optimized evidence theory.

[0073] Use different methods to measure the importance of nodes in the six networks of this data set, and sort the nodes in descending order based on their importance. Figure 3 lists the labels of the top 20 important nodes.

[0074] (I) Robustness and Vulnerability Evaluation

[0075] By plotting the "topN - G(N)" curve, analyze the impact of node removal, and evaluate the method in this chapter from the aspects of the robustness and vulnerability of the network. G represents the maximum connectivity coefficient of the network, and the expression is as follows:

[0076]

[0077] In the formula, N V is the total number of nodes in the network, and N R is the number of nodes belonging to the giant component in the network. If the value of G decreases more significantly as important nodes in the network are removed, the higher the accuracy of the node importance measurement method adopted.

[0078] The following focuses on analyzing the change trend of the network's largest connected set after sequentially removing the topN nodes according to node importance under a static attack. The experimental results are as shown in Figure 4 (a)-(f).

[0079] Figure 4The abscissa represents removing the top N nodes from the network, and the ordinate represents the maximum connectivity coefficient of the network. The experimental results show that in 6 different networks, with the removal of nodes, the MFF method shows a continuous downward trend in the maximum connectivity coefficient of the network compared with other methods; and after removing the same number of nodes, the maximum connectivity coefficient of the MFF method is always the minimum. This indicates that the important nodes measured by the MFF method are more accurate and have a greater impact on the network structure, that is, with the removal of the important nodes measured by the MFF method, the network structure is most affected. Especially in large-scale networks, the impact of removing important nodes on the network structure is significant. Thus, it can be concluded that when measuring the importance of nodes, in addition to considering node centrality and transitivity, the node reputation should also be considered, and the fusion method of each attribute is also crucial. Combining with the experimental result graph, it can be proved that the MFF method measures the importance of nodes more accurately compared with similar methods.

[0080] (2) SIR Propagation Characteristics Evaluation

[0081] By comparing the correlation between the propagation ability of nodes in the SIR model and various measurement values in 6 different networks, the performance of each method is evaluated. Since the network scales are different, the infection probability and recovery probability of the SIR model depend on the network topology. Generally, the setting of the infection probability depends on the asymptotic propagation threshold β of the network. The asymptotic propagation threshold is the minimum propagation probability that should be taken during the propagation process and is defined as:

[0082]

[0083] where <k>denotes the average degree of nodes in the network, <K 2 > denotes the second moment of the node degree, λ i is the amplification factor.

[0084] The recovery probability usually depends on the average degree of the network. The recovery probability γ is set as:

[0085]

[0086] where λ r is the amplification factor.

[0087] In the SIR model, there are generally two criteria for evaluating the importance of nodes: One is the average propagation range of the node, that is, the sum of the number of infected states I(t) and recovered states R(t) at the steady state. The larger the number, the more important the node. The other is the propagation rate, which can be evaluated from two aspects: the sum of the number of infected states I(t) and recovered states R(t) at the initial stage of propagation and the time taken to reach the steady state. The larger the number and the smaller the time taken, the more important the node; for the convenience of explanation, the sum of the number of infected and recovered states at the t-th time iteration step is denoted as F(t), and the results are as Figure 5 shown in (a)-(f) in

[0088] On the 6 networks, whether it is the propagation time or the propagation range, the MFF method can show better results compared with other methods, especially showing higher performance in the Advogato network. It is obvious that the performance of the MEC method is the worst. In any network, the MEVC and MFF methods are better than the MEC method, which shows that prestige has an important influence in determining the importance of nodes. Secondly, the performance of the MEVC method is inferior to that of the MEF method because the improved D-S evidence theory plays a key role in fusing centrality, transitivity, and prestige indicators, improving the accuracy of discriminating important nodes.

[0089] 【1】Kang Bingyi, Li Ya, Deng Yong, Zhang Yajuan, Deng Xinyang. Generation method and application of basic probability assignment based on interval numbers [J]. Acta Electronica Sinica, 2012, 40(06): 1092-1096.< / k>

Claims

1. A method for measuring the importance of social network nodes based on multi - feature fusion, characterized in that: Including the following steps: 1) First, preprocess the original feature data: Use the probability assignment generation method based on the interval number model to convert the feature values into the original basic probability assignment function BPA; Then, based on the evidence conflict and fuzziness, weights are assigned to the evidence and the original BPA is corrected. The method includes: calculating the centrality index, transitivity index, and prestige index of the nodes respectively; , are respectively expressed as the important probability and unimportant probability of node v i under the centrality index criterion; , are respectively expressed as the important probability and unimportant probability of node v i under the transitivity index criterion; , are respectively expressed as the important probability and unimportant probability of node v i under the prestige index criterion; Based on the evidence conflict degree and fuzziness, weights are assigned to the evidence, and the weights are used to correct the important and unimportant probabilities of node v i under the centrality index criterion, the important and unimportant probabilities of node v i under the transitivity index criterion, and the important and unimportant probabilities of node v i under the prestige index criterion; 2) In the evidence fusion stage, the BPA functions under centrality, transitivity, and prestige are used to perform evidence fusion with a new combination rule to obtain three new BPA functions. The corrected , , , , , are used to perform evidence fusion with the Dempster combination rule; Node v i Important probability , Node v i Unimportant probability , Node v i Probability of being important or unimportant The calculation formula is as follows: Among them: The rule for evidence fusion is to fuse the corrected , , into the importance probability of node v i , and fuse the corrected , , into the unimportance probability of node v i . The importance or unimportance probability of node v i is obtained by subtracting the importance probability of node v i and the unimportance probability of node v i from 1; 3) Obtain node v i of the multi - feature fusion node importance index value. The method includes assigning the probability of node vi being important or unimportant to the importance probability i of node v and the unimportance probability i of v . In the case of no prior knowledge, according to Bayesian theory, assign this probability to , to further correct the probability of node being important or unimportant. The correction method is as follows: Calculated is the basis for finally measuring the importance of nodes. After that, according to the values, the nodes in the social network are sorted in descending order. The nodes with higher rankings are more important.

2. The method for measuring the importance of social network nodes based on multi - feature fusion according to claim 1, characterized in that: The centrality index is obtained by dividing the out-degree of the node by the total number of nodes minus 1.

3. The method for measuring the importance of social network nodes based on multi - feature fusion according to claim 1, characterized in that: The transitivity index is represented by the linear negative correlation function of the local clustering coefficient of the node.

4. The method for measuring the importance of social network nodes based on multi - feature fusion according to claim 1, characterized in that: The prestige index is obtained by dividing the in-degree of the node by the total number of nodes minus 1.

5. The method for measuring the importance of social network nodes based on multi - feature fusion according to claim 1, characterized in that: , Constructed based on the probability assignment function using the centrality index, the maximum value of the centrality index, the minimum value of the centrality index, and the modification parameter λ, , Constructed based on the probability assignment function using the transitivity index, the maximum value of the transitivity index, the minimum value of the transitivity index, and the modification parameter δ, , Constructed based on the probability assignment function using the prestige index, the maximum value of the prestige index, the minimum value of the prestige index, and the modification parameter μ.