A Seed Expansion Community Detection Method and System Based on Community Clustering Features
By considering the characteristics and modularity increment of the community clustering in the seed expansion and community optimization stages, the problems of not considering the characteristics of community clustering and screening sparse communities in the prior art are solved, and the accuracy and recognition ability of community detection are improved.
Patent Information
- Application Number
- CN202210695902.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-20
AI Technical Summary
The existing community detection method based on seed expansion does not consider the clustering characteristics of the community, and there is a problem of inaccurate screening of sparse communities during the community optimization process.
A seed expansion club detection method based on the clustering characteristics of the society is proposed. The seed nodes are selected according to the clustering characteristics of the society in the seed expansion stage through the core community expansion algorithm, and the modularity increment is introduced in the community optimization stage to maximize the modularity.
It improves the accuracy of community testing, avoids the problem of inaccurate screening of sparse communities, and enhances the ability to identify complex online community structures.
Smart Images

Figure CN115018663B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network community discovery, and in particular, relates to a seed expansion community detection method and system based on community clustering characteristics. Background Art
[0002] The community detection method based on seed expansion first selects seed nodes as the initial community, and then expands the community according to the greedy strategy. The selection of seed nodes in the community detection method based on seed expansion is related to the community optimization strategy. There are various methods for selecting seeds. Cheng et al. select the node with the largest degree in the complex network as the seed node; Moradi et al. assign similarity values to each node in the complex network through link prediction, and then use the graph coloring algorithm to enhance the seed nodes; the LFM algorithm randomly selects seed nodes from the complex network; Hu et al. propose a voting-based seed node selection algorithm, and regard the node with the most votes as the seed node in each iteration; Whang et al. propose a strategy for selecting seed nodes based on the distance from the node to the clustering center; Liu et al. propose a strategy for selecting the node with the largest local clustering coefficient in the complex network as the seed node; Zhang et al. propose a seed node selection strategy based on random walk.
[0003] The concept of modularity was proposed by Newman et al. Modularity is usually used to evaluate the quality of community division in complex networks. The modularity optimization algorithm detects communities by maximizing modularity, and the process of community detection is also a process of modularity change. There are many optimization strategies for the modularity optimization algorithm: greedy algorithm, extreme optimization, spectral clustering, and simulated annealing (SA), etc. The greedy strategy is the most commonly used strategy. For example, the FN algorithm, the CNM algorithm, and the Louvain algorithm all adopt the greedy strategy. The FN algorithm uses the greedy strategy to select the community with the largest modularity increment or the least decrease to merge communities, and the community division corresponding to the maximum modularity is the community detection result of the algorithm. The execution process of the CNM algorithm is the same as that of the FN algorithm and also adopts the greedy strategy. Different from the FN algorithm, the CNM algorithm introduces data structures such as the modularity increment matrix in the algorithm to reduce the time complexity of the algorithm. The Louvain algorithm initializes each node in the complex network as a separate community, and then merges each node with its adjacent node with the largest modularity increment.
[0004] In the existing community detection method based on seed expansion, the clustering characteristics of the community are not considered during the seed expansion process, and there is a problem of inaccurate screening of sparse communities during the community optimization process. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a seed expansion community detection method and system based on community clustering features. Aiming at the problem of not considering community clustering features, the present invention expands communities with the clustering features of communities during the seed expansion process; aiming at the problem of inaccurate screening of sparse communities during the community optimization process, the present invention introduces modularity increment and merges communities with modularity maximization during community optimization, further improving the accuracy of community detection.
[0006] On the one hand, to achieve the above object, the present invention provides a seed expansion community detection method based on community clustering features, including the following steps:
[0007] Based on the clustering features of communities, use the core community expansion algorithm to obtain expanded communities;
[0008] Use the community optimization algorithm based on modularity to optimize the expanded communities and complete the community detection of complex networks.
[0009] Optionally, the method for obtaining expanded communities includes:
[0010] Based on the clustering features of communities, obtain the local clustering coefficient of nodes;
[0011] Based on the local clustering coefficient, obtain the community clustering coefficient for evaluating the community clustering ability;
[0012] Based on the community clustering coefficient and the degree of nodes, obtain seed nodes and sort the adjacent nodes in the seed nodes;
[0013] Calculate the clustering increment for the sorted adjacent nodes;
[0014] Based on the clustering increment, obtain expanded communities.
[0015] Optionally, the expression for calculating the clustering increment for the sorted adjacent nodes is:
[0016]
[0017] Wherein, represents the community clustering coefficient of community c when node v is included in the complex network t ; represents the community clustering coefficient of community c when node v and the edges associated with it are removed from the complex network t ; represents the clustering increment of node v for community c t .
[0018] Optionally, the method for obtaining expanded communities based on the clustering increment is:
[0019] Expand the nodes with the clustering increment greater than 0 to the seed nodes as the expanded communities.
[0020] Optionally, the method for optimizing the expanded communities by using the community optimization algorithm based on modularity is as follows:
[0021] Step 1: Establish a modularity increment matrix;
[0022] Step 2: Merge the communities with the modularity increment equal to the threshold in the modularity increment matrix;
[0023] Step 3: Repeat Step 1 and Step 2 until the values in the modularity increment matrix are all less than 0, and the optimization of the expanded communities is completed.
[0024] Optionally, the modularity increment matrix is:
[0025]
[0026] Wherein, represents the modularity of the entire complex network community division C corresponding to the communities c i and c j before merging, represents the modularity of the community result C of the entire complex network corresponding to the communities c i and c j after merging.
[0027] On the other hand, to achieve the above object, the present invention provides a seed-expanded community detection system based on community clustering characteristics, including: a seed expansion module and a community optimization module;
[0028] The seed expansion module is used to obtain the expanded communities by using the core community expansion algorithm based on the clustering characteristics of the communities;
[0029] The community optimization module is used to optimize the expanded communities by using the community optimization algorithm based on modularity, and complete the community detection of the complex network.
[0030] Optionally, the seed expansion module includes: a first obtaining unit, a second obtaining unit, a third obtaining unit, a calculation unit, and a fourth obtaining unit;
[0031] The first obtaining unit is used to obtain the local clustering coefficient of the nodes based on the clustering characteristics of the communities;
[0032] The second obtaining unit is used to obtain the community clustering coefficient for evaluating the community clustering ability based on the local clustering coefficient;
[0033] The third obtaining unit is configured to obtain seed nodes based on the community clustering coefficient and the degree of nodes, and sort the neighboring nodes in the seed nodes;
[0034] The calculating unit is configured to calculate the clustering increment for the sorted neighboring nodes;
[0035] The fourth obtaining unit is configured to obtain an extended community based on the clustering increment.
[0036] Optionally, the community optimization module includes: a construction unit, a merging unit, and an optimization unit;
[0037] The construction unit is configured to establish a modularity increment matrix;
[0038] The merging unit is configured to merge the communities with modularity increments equal to the threshold in the modularity increment matrix;
[0039] The optimization unit is configured to repeat the construction unit and the merging unit until the values in the modularity increment matrix are all less than 0, thereby completing the optimization of the extended community.
[0040] Compared with the prior art, the present invention has the following advantages and technical effects:
[0041] The present invention provides a method and system for detecting seed extended communities based on community clustering characteristics. In the seed extension stage, nodes with the largest local clustering coefficient and smaller degree are selected as seed nodes according to the clustering characteristics of the community, and then the community is extended with the clustering increment. In the community optimization stage, the present invention redefines the modularity increment function and merges communities with modularity maximization, avoiding the problem of inaccurate screening of sparse communities and further improving the accuracy of community detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0043] Figure 1 is a schematic flowchart of a method for detecting seed extended communities based on community clustering characteristics according to Embodiment 1 of the present invention;
[0044] Figure 2 is a schematic diagram of an undirected simple network according to Embodiment 1 of the present invention;
[0045] Figure 3 is a schematic diagram of the execution process of the SEA algorithm according to Embodiment 1 of the present invention;
[0046] Figure 4Schematic diagram of the experimental results of the LFR1 network in the first embodiment of the present invention. Among them, Fig. (a) is a schematic diagram of the NMI value of the LFR1 network experiment, and Fig. (b) is a schematic diagram of the Q value of the LFR1 network experiment;
[0047] Figure 5 Schematic diagram of the experimental results of the LFR2 network in the first embodiment of the present invention. Among them, Fig. (a) is a schematic diagram of the NMI value of the LFR2 network experiment, and Fig. (b) is a schematic diagram of the Q value of the LFR2 network experiment;
[0048] Figure 6 Schematic diagram of the experimental results of the LFR3 network in the first embodiment of the present invention. Among them, Fig. (a) is a schematic diagram of the NMI value of the LFR3 network experiment, and Fig. (b) is a schematic diagram of the Q value of the LFR3 network experiment;
[0049] Figure 7 Schematic diagram of the experimental results of the LFR4 network in the first embodiment of the present invention. Among them, Fig. (a) is a schematic diagram of the NMI value of the LFR4 network experiment, and Fig. (b) is a schematic diagram of the Q value of the LFR4 network experiment;
[0050] Figure 8 Schematic diagram of the comparison results of the SEA algorithm and the comparative algorithm in the first embodiment of the present invention on a real network;
[0051] Figure 9 Schematic diagram of the community detection results of the SEA algorithm in the first embodiment of the present invention on the Karate network;
[0052] Figure 10 Schematic diagram of the community detection results of the SEA algorithm in the first embodiment of the present invention on the Dolphins network;
[0053] Figure 11 Schematic diagram of the community detection results of the SEA algorithm in the first embodiment of the present invention on the Football network;
[0054] Figure 12 Schematic diagram of the community detection results of the SEA algorithm in the first embodiment of the present invention on the Riskmap network;
[0055] Figure 13 Schematic diagram of the structure of a seed expansion community detection system based on community clustering features in the second embodiment of the present invention. Detailed implementation manners
[0056] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail the present application.
[0057] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0058] Embodiment 1
[0059] As Figure 1 shown, the present invention provides a seed expansion community detection method based on community clustering features, including the following steps:
[0060] Based on the clustering features of the community, use the core community expansion algorithm to obtain the expanded community;
[0061] Use the community optimization algorithm based on modularity to optimize the expanded community and complete the community detection of the complex network.
[0062] In this embodiment, the community clustering coefficient, clustering increment, and modularity increment are defined. The community clustering coefficient is used to evaluate the clustering ability of a group of nodes in the complex network; the clustering increment is used to expand the core area of the community; the modularity increment is used to evaluate whether the modularity of the entire complex network community division can be improved before and after the merger of two communities.
[0063] In this embodiment, the sum average of the local clustering coefficients of the nodes in the community is called the community clustering coefficient. Assume C = {c1, c2,......, c n} represents the set of communities in the complex network. The set of nodes in community c t is represented by N(c t ) = {v t1 , v t2 ,......, v tm}. The community clustering coefficient is shown in Formula 1.1.
[0064]
[0065] Among them, LCC is the local clustering coefficient of the node. By analogy, the local clustering coefficient of the node is used to evaluate the clustering ability of the node in the complex network, and the community clustering coefficient reflects the clustering ability of a group of nodes in the complex network. The larger the value of the community clustering coefficient, the stronger the clustering ability of the community, and vice versa.
[0066] In this embodiment, the clustering increment is proposed based on the community clustering coefficient. The clustering increment represents the difference between the community clustering coefficient of the community when the complex network contains node v and the community clustering coefficient of the community when node v and the edges associated with node v are removed from the complex network. The definition of the clustering increment is as follows.
[0067] Assume V represents the set of all vertices in the complex network, and node v belongs to community ct The clustering increment of t is defined as shown in Equation 1.2.
[0068]
[0069] Among them, represents the community clustering coefficient of community c when node v is included in the complex network. t The community clustering coefficient of represents the community clustering coefficient of community c when node v and the edges associated with it are removed from the complex network. t The community clustering coefficient of represents the clustering increment of node v for community c. t When it means that node v improves the clustering ability of the community, otherwise it reduces the clustering ability of the community.
[0070] In this embodiment, the FN algorithm introduces the concept of modularity increment ΔQ. e ij represents the ratio of the number of edges connecting communities i and j to the number of edges in the entire complex network, a i and a j respectively represent the ratios of the sum of the node degrees within communities i and j to the sum of the degrees in the entire complex network. The formula for ΔQ is as shown in Equation 1.3
[0071] ΔQ = 2(e ij - a i a j ) (1.3)
[0072] Analyzing Equation 1.3, it can be seen that the modularity increment ΔQ is calculated based on the topological information of two adjacent communities, ignoring the impact of the merger of the two communities on other communities. Therefore, this paper proposes to consider the modularity increment ΔQ' of all communities. ΔQ' refers to the difference between the modularity of the communities in the entire complex network after the community merger and the modularity of the communities in the entire complex network before the community merger. The definition of the modularity increment ΔQ' is as follows.
[0073] Assume that the result of community partitioning of the complex network is represented by C = {c1, c2,......, c k}. c i and c j respectively represent any two communities in the community result C. The definition of the modularity increment ΔQ' is as shown in Equation 1.4.
[0074]
[0075] Among them, represents the modularity corresponding to the community partitioning C of the entire complex network before the merger of communities c i and c j , represents community ci and c j The modularity corresponding to the community result C of the entire complex network after merging. ΔQ'>0 indicates that the merging of two communities increases the modularity of the community division of the entire complex network, and ΔQ'<0 indicates that the merging of two communities reduces the modularity of the community of the entire complex network.
[0076] In this embodiment, the SEA algorithm is divided into a seed expansion stage and a community optimization stage. In the seed expansion stage, the CEA algorithm (Core Community Expansion Algorithm) selects the node with the largest local clustering coefficient and the smallest degree among the nodes that have not been divided into communities as the seed node, and then sorts the community neighbor nodes according to the local clustering coefficient, and calculates the clustering increment of the neighbor nodes for the community in turn according to Formula 1.2, and then expands the nodes with a clustering increment greater than 0 to the community. When there are no eligible neighbor nodes, a new community is created. The goal of the CEA algorithm is to divide the core communities of all communities in the complex network. When all nodes are divided into communities, the CEA algorithm ends. In the community optimization stage, the MOA algorithm (Modularity-based Community Optimization Algorithm) establishes a modularity increment matrix according to Formula 1.4, and selects the community with the largest increment for merging, and then re-establishes the modularity increment matrix and merges the communities until the values of the increment matrix are all less than 0 and the algorithm ends. The input of the MOA algorithm is the result output by the CEA algorithm, and then the communities are merged with modularity optimization.
[0077] The execution process of this embodiment: As Figure 2 shown, there are 7 nodes, including two communities, which are C1 = {v1, v2, v3, v4} and C2 = {v5, v6, v7} respectively. The execution process of the SEA algorithm is as Figure 3 shown
[0078] The SEA algorithm first calculates the local clustering coefficient (LCC) of the nodes in the complex network, and the calculation result is as Figure 3 (a) shown. According to the principle of the SEA algorithm to select the node with the largest local clustering coefficient and the relatively smallest node degree among the nodes that have not been divided into communities, the v2 node is selected as the seed node, as Figure 3 (b) shown. After selecting v2 as the seed node, the community nodes are sorted according to the local clustering coefficient. Then, according to Formula 1.2, calculate the clustering increment ΔLCC of the neighbor nodes v1 and v4 of the v2 node with respect to the community c1. The clustering increment of the node v1 for the community c1 is The clustering increment of the node v4 for the community c1 is According to the principle that the clustering increment of the core community expansion is greater than 0, v1 and v4 are expanded to the core area of the community c1, as Figure 3 (c) shown. Calculate the clustering increment ΔLCC of the neighbor nodes v3 and v5 of the community c1. The clustering increment of the node v3 for the community c1 is The clustering increment of node v5 to community c1 is Therefore, the core area of community c1 does not include nodes v3 and v5, and the expansion of community c1 ends. The expansion result is as shown in Figure 3 (d). Similarly, the SEA algorithm selects v6 as the seed node from the nodes in the complex network that have not been partitioned into communities, as shown in Figure 3 (e). Calculate the clustering increments of nodes v7 and v5 through formula 1.2. Among them, the clustering increment of node v5 to community c2 is The clustering increment of node v7 to community c2 is Expand nodes v7 and v5 to the core area of community c2, and then end the iteration. The processing result is as shown in Figure 3 (f). Finally, select node v3 in the complex network that has not been assigned to a community as the seed node. There are no nodes in the complex network that have not been assigned to a community for node v3, so it forms a separate community c3, as shown in Figure 3 (g). In the community optimization stage of the SEA algorithm, community c3 is reallocated. Calculate the community modularity increment ΔQ′(c3,c1) = 0.0625, ΔQ′(c2,c1) = -0.055 according to formula 1.4. Therefore, node v3 belongs to community c1. The execution result of the SEA method is as shown in Figure 3 (h). The communities detected by the SEA algorithm are c1 = {v1, v2, v3, v4} and c2 = {v5, v6, v7}.
[0079] The time complexity and space complexity of the algorithm are two important aspects for evaluating the performance of the algorithm. Due to the rapid development of computer hardware, the space complexity of the algorithm can be satisfied. In this embodiment, only the time complexity is used to evaluate the algorithm performance. The time complexity analysis of the SEA algorithm is as follows:
[0080] Assume that the number of nodes in the complex network G=(V, E) is n, the number of edges is m, the maximum degree of nodes is d, and the number of communities in C_pre is k.
[0081] (1) The time complexity of the CEA algorithm for calculating the local clustering coefficient of each node in the complex network is O(nd). When expanding the core community, the time complexity of sorting adjacent nodes by the local clustering coefficient is O(d log2 d), and the time complexity of expanding adjacent nodes is O(d). In the most extreme case, the number of local optimizations is n, so the time complexity is O(nd(log2 d + 1)). Therefore, the time complexity of the CEA algorithm is O(n(d(log2d + 1)+d))
[0082] (2) The time complexity of the MOA algorithm for calculating the modularity increment between communities is O(k 2) The number of community mergers is executed at most k times. Therefore, the time complexity of the MOA algorithm is O(k 3 ) The time complexity of the SEA algorithm is the sum of the time complexity of the CEA algorithm and the time complexity of the MOA algorithm. The time complexity of the SEA algorithm is O(n(d(log2d + 1)+d))+O(k 3 ).
[0083] To verify the effectiveness of the SEA algorithm for detecting communities proposed in the present invention, the present invention conducts experiments on the LFR benchmark network and real network datasets. The parameter settings of the LFR benchmark network of the present invention are shown in Table 1.
[0084] Table 1
[0085]
[0086] As shown in Table 1, the node scales of these 4 groups of networks are 50, 500, 1000, and 2000 respectively. In these 4 groups of networks, the average node degree is set to 10 in this chapter. The maximum degree of the LFR1 network is set to 20, and the maximum node degree of the remaining LFR benchmark networks is set to 40. The present invention sets the minimum community scale of the LFR1 network to 5 and the maximum community scale to 20. The present invention sets the minimum community scale of the LFR2, LFR3, and LFR4 benchmark networks to 10 and the maximum community scale to 50. The present invention sets the range of the mixing parameter of the four groups of LFR benchmark networks to be between 0.1 and 0.8.
[0087] The present invention uses modularity Q and normalized mutual information NMI as experimental evaluation indicators. The two evaluation indicators of modularity Q and normalized mutual information NMI have different focuses. Modularity Q evaluates the quality of community division of complex networks from the perspective of the topological structure of complex networks. NMI evaluates the accuracy of the algorithm for detecting communities from the consistency between the community division result and the real result.
[0088] The present invention selects six comparison algorithms: LPA, WalkTrap, Springlass, GN, Louvain algorithm, and NSA algorithm. Since the LPA algorithm is unstable, the LPA algorithm is executed 100 times in the experiment, and the sum of the average values of Q and NMI is taken respectively. During the experiment, according to the experimental results of the NSA algorithm with the threshold value between 0 and 1, the effect is optimal when the threshold value is 0.01. Therefore, the sparsity threshold of the NSA algorithm is set to 0.01.
[0089] Figure 4 (a), Figure 4(b) shows the experimental results of the SEA algorithm and its comparison algorithms on the LFR1 network. It can be seen from the figure that when μ ≤ 0.2, the NMI and Q values corresponding to the community detection of the SEA algorithm are the same as those of the Walktrap algorithm, Louvain algorithm, LPA algorithm, GN algorithm, and NSA algorithm. This shows that on the LFR1 network, when μ ≤ 0.2, the SEA algorithm can effectively detect communities just like the comparison algorithms. When μ > 0.2, the NMI and Q values of the community detection of the SEA algorithm are rapidly decreasing. The reason is that as the mixing parameter μ increases, the clustering characteristics of the communities in the LFR1 network also decrease, and the effect of the seed expansion of the SEA algorithm also decreases. From Figure 4 (a), Figure 4 (b), it can be seen that the SEA algorithm is more sensitive to the change of the μ value than the NSA algorithm. The SEA algorithm expands communities based on the community clustering characteristics, while the NSA algorithm expands communities based on node similarity.
[0090] Figure 5 (a), Figure 5 (b) shows the experimental results of the SEA algorithm and its comparison algorithms on the LFR2 network. It can be seen from the figure that when μ ≤ 0.3, the NMI and Q values corresponding to the community detection of the SEA algorithm are optimal compared with those of the Walktrap algorithm and Louvain algorithm. This shows that on the LFR2 network, when μ ≤ 0.3, the SEA algorithm can effectively partition communities from the LFR2 network. The NMI and Q of the community detection of the SEA algorithm are better than those of the NSA algorithm, indicating that the SEA algorithm further improves the accuracy of the seed expansion algorithm.
[0091] Overall, as the μ value increases, the accuracy of the community detection of the SEA algorithm decreases. The reason is that as the μ value increases, the community clustering characteristics of the LFR2 network also decrease.
[0092] Figure 6 (a), Figure 6 (b) shows the experimental results of the SEA algorithm on the LFR3 network. It can be seen from the figure that when μ ≤ 0.3, the NMI and Q values of the community detection of the SEA algorithm are slightly lower than those of the GN algorithm, Walktrap algorithm, and Louvain algorithm. It can be seen from the figure that the NMI and Q values of the community detection of the SEA algorithm are better than those of the NSA algorithm, which shows that on the LFR3 network, the SEA algorithm further improves the accuracy of the seed expansion community detection.
[0093] Figure 7 (a), Figure 7(b) shows the experimental results of the SEA algorithm and its comparison algorithms on the LFR4 network. It can be seen from the figure that the NMI and Q values of community detection by the SEA algorithm are slightly lower than those of the GN algorithm, Louvain algorithm, and Walktrap algorithm. It can be seen from the figure that the effect of community detection by the SEA algorithm is better than that of the NSA algorithm. This can prove that the SEA algorithm can improve the effect of community detection of the seed expansion algorithm.
[0094] Through the comparative experiments on four groups of artificial networks, namely LFR1, LFR2, LFR3, and LFR4, it can be known that as the mixing parameter increases, the accuracy of community detection by the SEA algorithm will decrease. Secondly, the experimental results show that the SEA algorithm can improve the accuracy of community detection of seed expansion.
[0095] The present invention conducts comparative experiments using four groups of real networks, namely Karate, Dolphins, Football, and Riskmap. All these four real networks have reference communities, and the present invention uses NMI and Q as evaluation indicators. The NMI and Q values of the SEA algorithm and the comparison algorithms on the real networks are shown in Table 2. The comprehensive results of NMI and Q of the SEA algorithm and the comparison algorithms on the real network dataset are as Figure 8 shown.
[0096] Table 2
[0097]
[0098] According to Table 2 and Figure 8 it can be known that on the Karate network, the NMI value corresponding to the community division by the SEA algorithm is lower than that of the Springlass algorithm, but better than that of the NSA algorithm and other comparison algorithms. Both the SEA algorithm and the NSA algorithm are based on seed expansion. The SEA algorithm is based on the clustering characteristics of communities. From the experimental results, the NMI of community detection by the SEA algorithm is better than that of the NSA algorithm. Similarly, the modularity Q corresponding to the community division by the SEA algorithm is higher than that of the NSA algorithm. From Figure 8 it can be known that the comprehensive values of NMI and Q of the community division by the SEA algorithm are larger than those of the community division by the NSA algorithm. Therefore, we can know that compared with the NSA algorithm, the SEA algorithm can more effectively detect communities from the Karate network. On the Dolphins network, the NMI and Q values of community detection by the SEA algorithm are the same as those of the Springlass algorithm, and better than other algorithms. On the Dolphins network, the NMI and Q values of community detection by the SEA algorithm are both better than those of the NSA algorithm. It can be seen that the SEA algorithm is effective on the Dolphins network, and at the same time, the SEA algorithm can improve the accuracy of community detection of the seed expansion algorithm. From Figure 8It can be seen that the comprehensive values of NMI and Q for community detection of the SEA algorithm on the Dolphins network are better than those of the comparative algorithms. On the Football network, the NMI and Q values corresponding to community detection of the SEA algorithm and the NSA algorithm are the highest, indicating that the SEA algorithm based on community clustering features and the NSA algorithm based on similarity are equally effective on the Football network. On the Riskmap network, the NMI of community detection of the SEA algorithm is slightly lower than that of the GN algorithm, but both the NMI and Q of community detection of the SEA algorithm are higher than those of the NSA algorithm, indicating that the method based on community clustering feature extension of the SEA algorithm is superior to the NSA algorithm on the Riskmap network. From Figure 8 it can be seen that the comprehensive values of NMI and Q corresponding to community detection of the SEA algorithm on the Riskmap network are slightly lower than those of the GN algorithm, but are better than those of other comparative algorithms.
[0099] To sum up, the community detection of the SEA algorithm on the four real networks is effective. Secondly, the SEA algorithm can further improve the accuracy of community detection of the seed expansion algorithm.
[0100] As Figure 9 shows the results of community detection of the SEA algorithm on the Karate network. The Karate network has 2 reference communities, and the SEA algorithm divides the Karate network into 4 communities. In the reference communities, the orange community where node 5 is located and the green community where node 0 is located belong to the same community. Analyzing the execution process of the SEA algorithm, we know that the algorithm divides nodes 16, 5, and 6 into the core community in the seed expansion stage. In the community optimization stage, the nodes that can improve the modularity of the entire network community are merged. Since the merger of the orange community where node 5 is located and the green community where node 0 is located cannot improve the modularity of the community, they are divided into two communities.
[0101] Figure 10 shows the results of community division of the SEA algorithm on the Dolphins network. The Dolphins network has 2 reference communities, and the SEA algorithm divides the Dolphins network into two communities. As can be seen from Table 3-1, the NMI of community division of the SEA algorithm is 0.5865, and the Q value is 0.5285. Although the effectiveness of community division of the SEA algorithm on the Dolphins network is better than that of the comparative algorithms, the result of community division is not ideal. Observing Figure 10 it can be seen that the scale of the red community is large and the scale of the green community is small. The reason is that the SEA algorithm divides a group of nodes with strong aggregation in the local network into a core community in the seed expansion stage, and merges them into the red community with the goal of maximizing the modularity of the entire community division in the community optimization stage of the algorithm.
[0102] Figure 11It shows the community detection results of the SEA algorithm on the Football network. There are 12 reference communities in the Football network, and the SEA algorithm divides the Football network into 10 communities. Among the reference communities, nodes 93, 90, and 50 belong to three different communities respectively, but the SEA algorithm divides them into one community. The reason is that when merging communities, the SEA algorithm maximizes modularity and merges the three communities where these three nodes are located into one large community.
[0103] Figure 12 It shows the community detection results of the SEA algorithm on the Riskmap network. There are 6 reference communities in the Riskmap network, and the SEA algorithm divides the Riskmap network into 7 communities. In the reference communities, the orange community where node 17 is located and the blue community where node 22 is located belong to the same community except for node 12. However, the SEA algorithm divides them into two communities. The reason is that during the seed expansion stage of the SEA algorithm, they are divided into two core communities, and during the community optimization stage, since the merger of the two communities cannot improve the modularity of the community, they are divided into two communities.
[0104] Apply this embodiment to the extraction of protein complexes in the protein-protein interaction network. For example, in the protein-protein interaction network, nodes represent proteins, and edges represent the interaction relationship between proteins. Research shows that proteins generally do not act alone, and protein complexes composed of multiple proteins have biological significance. The protein-protein interaction network has the characteristics of community structure, and each community in the community structure corresponds to a protein complex. This technical solution proposes a two-stage protein complex extraction scheme based on the clustering characteristics of communities in the protein-protein interaction network. The preliminary extraction of protein complexes is completed in the seed expansion stage; the optimization of protein complexes is completed in the community optimization stage based on the preliminary extraction of proteins. This scheme defines the community clustering coefficient for evaluating communities according to the clustering characteristics of communities, and defines the community clustering increment based on the community clustering coefficient. In the seed expansion stage, select the node with the largest local clustering coefficient and relatively small degree in the protein-protein interaction network as the seed node. Then, sort the neighboring nodes of the community according to the local clustering coefficient, and expand the neighboring nodes with a clustering increment greater than 0 into the community. In the community optimization stage, the SEA algorithm defines the modularity increment and introduces the modularity increment matrix during the optimization process. Each time during iteration, select the two protein complexes with the largest modularity increment for merging.
[0105] Embodiment 2
[0106] As Figure 13 shown, the present invention also provides a seed expansion community detection system based on community clustering characteristics, including: a seed expansion module and a community optimization module;
[0107] The seed expansion module is used to obtain an extended community by using the core community expansion algorithm based on the clustering characteristics of the community.
[0108] The community optimization module is used to optimize the extended community by using the community optimization algorithm based on modularity, and complete the community detection of the complex network.
[0109] Furthermore, the seed expansion module includes: a first acquisition unit, a second acquisition unit, a third acquisition unit, a calculation unit, and a fourth acquisition unit;
[0110] The first acquisition unit is used to obtain the local clustering coefficient of a node based on the clustering characteristics of the community.
[0111] The second acquisition unit is used to obtain the community clustering coefficient for evaluating the community clustering ability based on the local clustering coefficient.
[0112] The third acquisition unit is used to obtain seed nodes based on the community clustering coefficient and the degree of the node, and sort the adjacent nodes in the seed nodes;
[0113] The calculation unit is used to calculate the clustering increment for the sorted adjacent nodes;
[0114] The fourth acquisition unit is used to obtain the extended community based on the clustering increment.
[0115] Furthermore, the community optimization module includes: a construction unit, a merging unit, and an optimization unit;
[0116] The construction unit is used to establish a modularity increment matrix;
[0117] The merging unit is used to merge the communities with modularity increments equal to the threshold in the modularity increment matrix;
[0118] The optimization unit is used to repeat the construction unit and the merging unit until the values in the modularity increment matrix are all less than 0, and complete the optimization of the extended community.
[0119] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A seed expansion community detection method based on community clustering features, which is applied to the extraction of protein complexes in a protein interaction network, and is characterized in that, It includes the following steps: Based on the clustering characteristics of communities, using the core community expansion algorithm, an extended community is obtained, and each community in the community structure corresponds to a protein complex; Using the community optimization algorithm based on modularity, optimize the extended community to complete the community detection of the complex network; The sum average of the local clustering coefficients of the nodes in a community is called the community clustering coefficient. Suppose \(C = \) represents the set of communities in a complex network. The community node set is denoted by \(N( \) ) = The community clustering coefficient is: Among them, LCC is the local clustering coefficient of the node. The local clustering coefficient of the analogous node is used to evaluate the clustering ability of the nodes in the complex network. The community clustering coefficient reflects the clustering ability of a group of nodes in the complex network. The larger the value of the community clustering coefficient, the stronger the clustering ability of the community, and vice versa; The method for obtaining the extended community includes: Based on the clustering characteristics of communities, obtain the local clustering coefficient of nodes; Based on the local clustering coefficient, obtain the community clustering coefficient for evaluating the community clustering ability Coefficient; Based on the community clustering coefficient and the degree of nodes, obtain seed nodes, where nodes represent proteins, and sort the adjacent nodes in the seed nodes; Calculate the clustering increment for the sorted adjacent nodes; Based on the clustering increment, obtain the extended community; The expression for calculating the clustering increment for the sorted adjacent nodes is: , Among them, represents the community clustering coefficient of the community when the complex network contains node ; represents the community clustering coefficient of the community when removing node and its associated edges from the complex network. The edge indicates that there is an interaction relationship between proteins. represents the clustering increment of node for the community ; The method for obtaining the extended community based on the clustering increment is: Expand the nodes with the clustering increment greater than 0 to the seed nodes as the extended community; The method for optimizing the extended community using the community optimization algorithm based on modularity is: Step 1: Introduce a modularity increment matrix during the optimization process, and select the two protein complexes with the largest modularity increment to merge each iteration; Step 2: Merge the communities with the modularity increment equal to the threshold in the modularity increment matrix; Step 3: Repeat Step 1 and Step 2 until the values of the modularity increment matrix are all less than 0, and complete the optimization of the extended community; The modularity increment matrix is: , Among them, represents the community and the modularity corresponding to the community division C of the entire complex network before merging, represents the community and the modularity corresponding to the community result C of the entire complex network after merging.
2. A seed expansion community detection system based on community clustering features, which is applied to the extraction of protein complexes in a protein interaction network, and is characterized in that, It includes: A seed expansion module and a community optimization module; The seed expansion module is used to obtain an extended community based on the clustering characteristics of communities using the core community expansion algorithm, and each community in the community structure corresponds to a protein complex; The community optimization module is used to optimize the extended community using the community optimization algorithm based on modularity to complete the community detection of the complex network; The seed expansion module includes: a first obtaining unit, a second obtaining unit, a third obtaining unit, a calculation unit, and a fourth obtaining unit; The first obtaining unit is used to obtain the local clustering coefficient of nodes based on the clustering characteristics of communities; The second obtaining unit is used to obtain the community clustering coefficient for evaluating the community clustering Ability; The third obtaining unit is used to obtain Seed nodes, where nodes represent proteins, and sort the adjacent nodes in the seed nodes; The calculation unit is used to calculate the clustering increment for the sorted adjacent nodes; The fourth obtaining unit is used to obtain the extended community based on the clustering increment; The community optimization module includes: a construction unit, a merging unit, and an optimization unit; The construction unit is used to establish a modularity increment matrix; The merging unit is used to merge the communities with the modularity increment equal to the threshold in the modularity increment matrix; The optimization unit is used to repeat the construction unit and the merging unit until the values of the modularity increment matrix are all less than 0, and complete the optimization of the extended community; The sum average of the local clustering coefficients of the nodes in a community is called the community clustering coefficient. Suppose \(C = represents the set of communities in a complex network. The community node set is represented by \(N( ) = . The community clustering coefficient is: Among them, LCC is the local clustering coefficient of the node. The local clustering coefficient of the analog node is used to evaluate the clustering ability of the nodes in the complex network. The community clustering coefficient reflects the clustering ability of a group of nodes in the complex network. The larger the value of the community clustering coefficient, the stronger the clustering ability of the community, and vice versa; The expression for calculating the clustering increment for the sorted adjacent nodes is: , Among them, represents the community clustering coefficient of the community when the complex network contains node ; represents the community clustering coefficient of the community when the node and its associated edges are removed from the complex network. The edge indicates that there is an interaction relationship between proteins. represents the clustering increment of node for community ; The method for obtaining the extended community based on the clustering increment is as follows: Expand the nodes with the clustering increment greater than 0 to the seed nodes as the extended community; The process of optimizing the extended community using the community optimization algorithm based on modularity is as follows: Step 1: Introduce a modularity increment matrix in the optimization process. At each iteration, select the two protein complexes with the largest modularity increment and merge them; Step 2: Merge the communities with the modularity increment equal to the threshold in the modularity increment matrix; Step 3: Repeat Step 1 and Step 2 until the values in the modularity increment matrix are all less than 0, and the optimization of the extended community is completed; The modularity increment matrix is: , Among them, represents the community and the modularity corresponding to the community division C of the entire complex network before merging, represents the community and the modularity corresponding to the community result C of the entire complex network after merging.
Citation Information
Patent Citations
Method for discovering complex network fuzzy association based on membership transmission
CN104657418A